Limitations and objections
The questions we get asked, with the answers we would give in a seminar.
- Isn't a single year of data too noisy?
Yes, and it cuts in the paper's favour. Value-added is reported unshrunk, so one σ is one standard deviation of the estimated teacher-effect distribution, which is inflated by estimation error. Scaling effects by an inflated standard deviation makes the reported gradients conservative, not generous.
It also means the intervals are wide, which is why the route remainder is reported as a failure to detect rather than a null.
- The +0.60σ jump from first to second year — is that really learning?
We cannot say that it is. It compares two different cohorts measured in a single wave, so it cannot separate learning from cohort composition or from attrition of the teachers who left after year one.
- Does this show the programme works?
No. It evaluates the teachers the programme supplies, not the programme. The relevant comparison for a programme evaluation is the marginal hire that a vacancy would otherwise have drawn, and that teacher is never observed in this data.
- Is a near-zero route residual evidence that the routes are the same?
No. The 95% interval is [−0.37, +0.17], wide enough to contain a moderate shortfall. And the residual is a remainder: it absorbs training, motivation, unobserved selection and any ability-by-experience interaction at once.
- Could this be about which schools these teachers work in?
It could, and the design cannot fully rule it out. With about 1.9 sampled teachers per school, school fixed effects are not supported; what the data supports is a control for school faculty composition. The paper reports that adverse test rather than hiding it.