We ran our own test on 30 students. It failed.

In July we took a 24-question instrument into a Tier-1 English-medium school and ran it on thirty students. The sheets were scanned and machine-read. Twenty-nine were fully scoreable.

The result was not a partial success. It was a zero.

What the numbers said

Internal consistency measures whether the questions on a scale agree with each other. If a set of questions is measuring one thing, a student who answers one way on the first should tend to answer the same way on the rest. Ours did not.

The mean correlation between items was −0.017. Of 276 possible item pairs, only seven correlated above 0.4. Every eigenvalue of the item correlation matrix fell below what random data would produce. There was no underlying factor at all — not the two the instrument assumed, not even one.

Broken out by block, the story got worse in an interesting way. The story-based situational items — the ones the whole design was built around — scored −0.45 and −0.22. The plain, context-free items written almost as an afterthought scored +0.19 and +0.31.

The part we were proud of performed worse than the part we nearly cut.

Why it happened

Three things, and none of them were about wording.

  • The setting supplied a correct answer. Every situational item was set at a school festival. Once a student pictures parents arriving and a hall to be set up, the question stops being "what would I rather do" and becomes "what does this situation need". That is a competence judgement, not a preference.
  • The setting supplied a hierarchy. The front desk is the senior job. Decorations are junior work. Prestige walked back in through the scene rather than the task — the exact thing the instrument exists to remove.
  • The setting supplied an audience. A teacher asked. Other volunteers were watching. Part of every answer was performance.

Separately, and probably compounding all three: comprehension. In conversation afterwards it became clear that students did not know the meaning of ordinary words in the items. One of them was "seek". Our readability check counted syllables. "Seek" is one syllable and passed every check we had. Those checks measure how easy a word is to pronounce, not whether a fifteen-year-old knows it.

What we did

We did not ship it.

That sounds obvious written down. It was not obvious at the time. The instrument produced numbers. The numbers could have been put in a report and the report could have been given to thirty students, and none of them would have known that what they were reading was noise.

Instead: the situational block was cut entirely. A reading check now runs before anything is scored, and a student who cannot read the questions is not placed — the report says so. The remaining items were rebuilt, and the vocabulary gate was rebuilt around word-list membership rather than syllable count.

Why we are writing this down

Career assessment in India is sold with a confidence the underlying science does not support. Instruments are not published, reliability figures are not shared, and every student who sits down gets a definitive answer regardless of whether the instrument could read them.

We would rather be the company that publishes the failure. An instrument that returns noise is not asking a student anything. It is the appearance of asking, which is worse than not asking, because it launders a guess as a finding.

← All writing