19 readers opened the test booklets for every grade-5 program and transcribed 121 vocabulary questions verbatim. Every count below is built on those transcriptions.
Verdict: Vocabulon, by a moderate margin.
4–0 judge panel, upheld by both adversarial reviews. The statistics review found two flaws in our own pipeline; correcting them widened Vocabulon's lead. The margin stays moderate because grade 5 is the close grade: several programs lean the other way, and about 45% of items reward skills no word list can teach.
121
real questions transcribed verbatim from 19 programs' booklets
2.2×
lead on explicitly-tested tier-2/3 words (27.5% vs 12.5%)
of real items are list-proof; they reward context-clue skill no list covers
Ground truth · Real questions, read
What grade-5 vocabulary questions look like in the booklets
Every card is a verbatim transcription. A blue edge means the item favors Vocabulon, orange means it favors VocabLoco, and gray means neither list can win it.
Florida FAST · Grade 5 · 2025isolate — Vocabulon g5 · not on VocabLoco
Read this sentence from the passage. "In fact, these beaches are so isolated you have to take a boat to them." What is the meaning of isolated as it is used in the sentence?
isolated is the most-retested target in the corpus, an explicit question in FAST, IAR, and PARCC, and one of only two grade-5 words tested across independent item banks. Vocabulon lists it at grade 5. VocabLoco omits it at every grade.
Colorado CMAS · Grade 5 · 2024rural — Vocabulon g4 · not on VocabLoco
Part A: What is the meaning of rural as it is used in paragraph 3 of the passage from The Renaissance? A appealing to the people B related to the country C dedicated to growth D full of opportunity
The hardest vocab item on that CMAS form: 36% of students got it right. A cold word-knowledge target, taught by Vocabulon a year before the test.
Illinois IAR / PARCC · Grade 5answer choices — Vocabulon words
Part A. What is the meaning of the word range as it is used in paragraph 3 of the article "Helping Giant Pandas"? A. territory B. continent C. zoo exhibit D. mountainous terrain
The target (range) is on neither list. The correct answer (territory) and two distractors (continent, terrain) are all Vocabulon g4–5 words, and none are VocabLoco's. Parsing the choices requires the challenger's vocabulary. Across 555 teachable tier-2/3 answer-choice words, Vocabulon covers 27.4% and VocabLoco 19.6%.
New York NYS · Grade 5 · 2008dignity — Vocabulon g5 · not on VocabLoco
"He wasn't hurt, except for his dignity—the sauce in his beautiful white feathers turned him splotchy orange for several weeks." Which word means about the same as "dignity"?
A pure synonym item on a tier-2 abstract noun. VocabLoco covered 0 of 8 transcribed NYS grade-5 targets.
Kentucky KSA · Grade 5 · 2023primary — VocabLoco g4 · not on Vocabulon
Which phrase would best replace "primary method" as it is used in paragraph 2? A easy system B main process C modern answer D creative solution
A VocabLoco win. It carries both primary and method plus most of the choice words, while Vocabulon misses primary entirely. Only 60% of Kentucky students answered it correctly.
Florida FCAT 2.0 · Grade 5 · 2012humble — VocabLoco g3 · not on Vocabulon
"But in the centuries following the humble yet beautiful career of 'the Backwoods Boy' from the hut to the White House…" Which word has the same meaning as humble?
A VocabLoco win. FCAT ran 2/7 for VocabLoco vs 0/7. FCAT, Kentucky, NC EOG, and RICAS all lean VocabLoco; check them against Alpha's test map.
New York NYS · Grade 5 · 2008list-proof · multiple-meaning
"A spitting cobra temporarily blinded him with a jet of venom." In this sentence, what does the word "jet" mean? A a dark color B a bright flash C a forceful spray D a hissing sound
An everyday word tested in a rare secondary sense. About 19 of the 121 items work this way (jet, mount, fold, picture=imagine). Only context-clue instruction wins these points.
Massachusetts MCAS · Grade 5 · 2025list-proof · sense discrimination
"…your More with Les paper keyboard is exactly to scale." What does the phrase "exactly to scale" mainly suggest about the paper keyboard?
This item shows why naive counting misleads. VocabLoco lists scale at g3 and again at g4, and it buys nothing here, because the distractors are the other senses of scale. Counting scale as covered credits VocabLoco for a point the list does not win.
"…said the Dewdrop as it mingled its tiny drop with the running river." What does the word mingled mean in this sentence? A joined B ran C talked D touched
mingled is on neither list, yet 72% of students answered correctly from context alone. Direct evidence that context-clue skill carries many of these items.
What the 121 transcribed items test
blue: styles a list can help with · gray: styles no list can
explicit vocab
37
figurative/idiom
25
evidence-paired
22
multi-meaning
19
affix/root
11
craft/meta
7
55 of 121 (~45%) are structurally list-proof: figurative phrases ("tip the scales", "a whale of a problem"), common words in secondary senses, or items that hand the student the root in the stem (FCAT prints "communicare, meaning to take part in"). Of the 121 targets, 31 are on Vocabulon, 27 on VocabLoco, and 76 on neither.
Exhibit A · Tested-word coverage
The words tests explicitly ask about
Census: 120 tier-2/3 grade-5 vocabulary-question targets across all programs, regex-mined and cross-checked against the corpus's own parser fields (the cross-check added 4 words and changed no shares). At grade 5, every grade 3–5 placement is taught by test day, so any-grade coverage is already the timing-correct metric.
VocabulonVocabLoco
All tested tier-2/3 targets · n=120
Vocabulon
27.5%
VocabLoco
12.5%
Strictly teachable at g5 (majority don't know it before grade 5) · n=80
Vocabulon
21.3%
VocabLoco
10.0%
STAAR grade 5, the named must-win program · 29 explicit targets
Vocabulon
7 of 29
VocabLoco
4 of 29
data table + adversarial recomputations
Metric
n
VocabLoco
Vocabulon
Tested tier-2/3, any grade
120
12.5%
27.5%
Teachable kg≥4
114
12.3%
28.1%
Strictly teachable kg≥5
80
10.0%
21.3%
After reviewer's family dedup (RICAS=MCAS etc.)
—
24.2%
51.5%
After reviewer's symmetric stemming
—
35.9%
56.2%
STAAR explicit targets (stemmed)
29
4 (5)
7 (8)
Combined census incl. parser fields
118
13.6%
28.0%
The adversarial recomputation changed the numbers here. The reviewer found the corpus dedup had missed rebranded programs (RICAS is MCAS rebadged; NY 2016–18 and KY 2021–22 overlap) and that strict matching undercounts both lists. After both errors were corrected, Vocabulon's lead was wider than before.
Exhibit B · Slots & quality
Grade-5 slot economics
VocabLoco · 272 grade-5 words167 sweet-spot slots
9%
25%
61% sweet spot
Vocabulon · 306 grade-5 words198 sweet-spot slots
6%
21%
65% sweet spot
9%
junk (tier 0/1)already known by g4sweet spot (kg 5–7)stretch (kg 8+)
VocabLoco's grade-5 page is far cleaner than its grade-3/4 pages (9% junk vs 34% at g4), so this is its best grade. Vocabulon's edge here comes from productive slots (198 vs 167), passage coverage (24.9% vs 18.1% of teachable passage words), and answer-choice vocabulary (27.4% vs 19.6%). One confound: Vocabulon's grade-5 page is also 12.5% larger, enough to nudge coverage figures but not to explain a 2.2× tested-target gap.
Self-audit · What we caught in our own work
Corrections from our own review
Retracted: the convergence table
We previously showed Vocabulon leading "words appearing in ≥K programs" at every K. After deduplicating to 13 true program families, K≥5 and K≥8 flip to VocabLoco (K≥8: 44.0% vs 28.0%, n=25). Convergence counts generic common passage words rather than tested targets, so it is excluded from the case in both directions. The tested-target metrics survive the same dedup and widen.
Corrected: item-bank double counting (also fixed on the grade-4 page)
"spectacular ×5" was one PARCC/IAR question duplicated across exports, and RICAS's tested words are a near-subset of MCAS's (Jaccard 0.83–0.97). All counts now collapse shared banks and count each unique word once. The grade-4 "3×" timing headline became 1.9× under exact recomputation.
Flagged: the raw item tally is the most VocabLoco-favorable count
On the 121 transcribed items the raw target tally is 31–27, nearly a tie. That tally double-counts the MCAS/RICAS pool and includes 55 list-proof items. Unique-word matching gives 36–19 strict and 39–28 stemmed. All three figures appear above; the verdict holds under the tally most favorable to the incumbent.
Known limits we can't fix from this corpus
Difficulty labels are model-rated (audited at 90% tier / 98% known-grade agreement, n=61), so gaps under two points should be read as ties. No single cut is significant alone, since STAAR is a 3-item margin, and the cuts share one pipeline. MAP items aren't public, so MAP similarity is argued by proxy. The expected score gain from switching lists alone is likely under one point per test form. The add-list and context-skill instruction carry more points than the switch itself.
Exhibit C · The patch
The gap, and the 139-word add-list
76 of 121 real item targets and 20 of 29 STAAR targets are on neither list. Closing the previously-tested gap takes 185 curated adds on Vocabulon vs 532 on VocabLoco. The 139 highest-priority adds, led by tested targets and the question-stem metalanguage that gates whole item types:
VocabLoco's wins come along:primary, humble, and transparent, the words behind the items VocabLoco won, are all in the add-list.
What would make this stronger
Open questions an approver should ask
Grade 6 is a hole in the product itself. The approval scope is G4–6, the corpus holds 631 grade-6 documents, and both lists stop at grade 5. A grade-6 list needs to exist before a G4–6 approval means anything.
NY item maps could grow the tested-word census tenfold. New York publishes per-item standards tags (L.5.4 = vocabulary) with item numbers. Joining them to released booklets would give an authoritative multi-year census; today's counts are consistent across three independent views but run on smaller n.
No bridge yet from coverage percentages to test-score points. Blueprints and answer keys would let us compute what share of raw points vocabulary items carry per program, which is what the literal "90%+" criterion needs.
A lemma-level recount (word families rather than strict inflections) should be a signoff condition. It recovers ~9 points for VocabLoco and ~5 for Vocabulon, flips nothing, and would remove a known undercount on both sides.
Check Alpha's test map against the programs that lean VocabLoco (FCAT, Kentucky, NC EOG).
Appendix · Method
How this was built
Reading first19 reader agents opened the actual booklets for every grade-5 program, transcribed 121 vocabulary items verbatim with answer choices, classified each item's style, and checked every target against both list CSVs.
Census295 grade-5 docs → 66 administrations (→ ~57 after the reviewer's family dedup) across 17 programs; 120 explicit tier-2/3 targets; 4,446 corpus words and both grade-5 pages fully tier/known-grade classified (audit: 90%/98% agreement).
New signal: answer-choice vocabulary1,165 items' options mined, yielding 555 teachable tier-2/3 choice words a student must know to parse the question at all.
Judging & reviewFour lenses (coverage, difficulty, guarantee, read-evidence) scored blind to which list is in-house. A statistics reviewer with full data access recomputed everything and found two real flaws, whose corrections widened the lead. A skeptical-approver review and a completeness critic round out the record; the critic's findings are the section above.