Grade 4 is the deepest slice of the corpus: 380 documents, 70 unique test administrations, 19 programs. This brief measures both lists against all of it.
"The right word list is the correct word list that allows them to deliver a 90%+ on any state or Common Core reading comprehension test… You just pull down tests and get all the state tests and see if your word list is going to teach all of those."
— Andy, Aug 27 demo (the approval criterion)
Verdict: Vocabulon, by a strong margin.
Unanimous 3–0 judge panel across coverage, difficulty, and guarantee lenses. The margin held through adversarial review. A second audit caught two inflation artifacts, item-bank duplication and a subsample ratio, and both corrections appear below.
3–0
lens panel, all three for Vocabulon (2 strong, 1 moderate)
1.9×
lead on tested words taught by test day (20.5% vs 11.1%)
0/2
adversarial refutations landed; both upheld the verdict at high confidence
4,686
grade-4 corpus words tier-classified to find every tier-2/3 target
Before the exhibits · The naive count
The quick spreadsheet check misleads in both directions
Counting raw word overlap produces three plausible-looking numbers. All three ignore the same question: does the word teach a grade-4 kid anything before test day?
Three naive counts and what's wrong with each
L = VocabLoco · U = Vocabulon
Naive metric
VocabLoco
Vocabulon
Why it misleads
Tested words on each product's grade-4 page only
7.7%
7.7%
A dead tie, because it ignores grade placement: grade-3 words are already taught by test day, grade-5 words are not
STAAR G4 explicit vocab targets covered (n=25)
4
4
Tie again, and 18 of 25 are on neither list. The STAAR story is the gap
List words found anywhere in test docs, frequency-weighted
9.9
6.3
VocabLoco "wins" here because because, although, and internet match text everywhere while teaching nothing. The slice is also inflated by test boilerplate
Even so, the simplest defensible count (raw any-grade overlap with tested words, no filters at all) still picks Vocabulon, 27.2% vs 16.8%. No counting method we could construct flips the winner. The exhibits below show why.
Exhibit A · What tests ask
Coverage of the words tests explicitly test
We regex-mined every "What is the meaning of ___?" style question across all 2,191 documents: 167 grade-4 vocabulary targets, of which 117 are teachable tier-2/3 words (words a majority of grade-4 students don't already know). Explicit targets are the strongest signal: these are the words test writers choose.
VocabulonVocabLoco
Tested tier-2/3 grade-4 words covered by each list
percent of tested words on the list · strict inflectional matching (isolate ⇒ isolated)
On the list at any grade 3–5 · n=117 teachable targets
Vocabulon
28.2%
VocabLoco
16.2%
Taught by test day: only grade ≤ 4 placements count, which is how school sequences the words
Vocabulon
20.5%
VocabLoco
11.1%
data table
Metric
n / weight
VocabLoco
Vocabulon
Any grade 3–5, all tier-2/3 tested (raw)
125
16.8%
27.2%
Any grade 3–5, teachable (kg≥4)
117
16.2%
28.2%
Timing-corrected (grade ≤4 placements)
117
11.1%
20.5%
Words covered (of 117 teachable)
117
19
33
The timing correction was the adversarial reviewers' strongest objection to the headline number: "any grade 3–5" credits grade-5 placements a grade-4 tester has never been taught. Recomputed exactly, Vocabulon holds 20.5% vs VocabLoco's 11.1%, still nearly 2:1, because six of VocabLoco's nineteen hits (abruptly, retreat, transformation…) sit only on its grade-5 page.
The 30 most-tested grade-4 words, and who has them
Ordered by how widely each word was tested. After collapsing shared item banks (IAR/PARCC/NJSLA, MCAS/RICAS, STAAR variants), nearly every word here is one unique test item, so the unit of counting is the word rather than the document. L/U shows which list carries the word and at what grade. The words with no tag on either list are the subject of Exhibit D.
Exhibit B · Where the slots go
How each list spends its teaching slots
Every one of the 634 grade-4 list words was classified by Beck tier and by the grade at which a majority of US students already know it (dual-rated for tested words, independently audited at 89% tier / 99% known-grade agreement). Here is what each list spends its slots on:
VocabLoco · 368 grade-4 words194 productive slots
34% junk
10%
53% sweet spot
Vocabulon · 266 grade-4 words195 productive slots
5%
73% sweet spot
17%
junk (tier 0/1)already known by g3sweet spot (kg 4–6)stretch (kg 7+)
The two lists carry an almost identical number of productive words, 194 and 195. VocabLoco needs 102 extra slots of a child's practice time to deliver the same payload. 44% of its grade-4 list teaches nothing a typical fourth-grader doesn't already have.
What's on VocabLoco's grade-4 list
A sample of the 125 tier-0/1 entries: sight words, function words, and spelling-drill items no ELA test asks the meaning of.
Vocabulon's own weak tail: 46 words (17.3%) sit at known-grade 7+, mostly social-studies curriculum vocabulary (feudalism, shogun, serf, ratify, dynasty, monastery). Some of that is justified stretch, since tests do target kg-6–8 words, and part is waste for ELA test prep. Total waste comes to roughly 27%, against VocabLoco's 44% or more.
Exhibit C · The passages
Passage vocabulary: a smaller edge
Beyond explicit vocab questions, kids meet hard words inside reading passages. Among the 900 teachable tier-2/3 words appearing in 3+ of the 70 unique grade-4 test administrations, Vocabulon covers 24.2% vs VocabLoco's 18.1%. The "convergence method" (words appearing across N independent test programs — the method already used for the grade-3 ITBS analysis) gives the same picture at every threshold:
Convergence coverage — tier-2/3 words appearing in ≥ K of 19 grade-4 test programs
≥ K programs
words
VocabLoco
Vocabulon
3
2,125
12.2%
13.9%
4
1,281
13.3%
15.1%
5
803
14.7%
16.4%
6
505
17.4%
18.4%
8
196
20.4%
19.9%
Near parity, with a slight Vocabulon edge. Grade 4 is not the grade-3 blowout (74% vs 34% on ITBS), which is why the slot-quality and tested-target evidence above carries the verdict.
Exhibit D · The gap neither list covers
Neither list gets near 90% today
67%
of the 117 teachable explicitly-tested grade-4 tier-2/3 words — including (spectacular, din, shrieks, applauded, composure) — are on neither list.
The "guarantee we cover everything previously tested" therefore has to be delivered by adding words. That turns the decision question into: which list is the better base to patch?
Vocabulon as chassis
Starts at 28.2% tested-word coverage
270 curated adds close the gap
Nothing to remove (5% junk)
Selection already tracks what tests target
VocabLoco as chassis
Starts at 16.2% tested-word coverage
324 adds needed, a bigger patch
Plus ~a third of the list should be pruned
Ends up rebuilt into a different product
The recommended add-list for Vocabulon grade 4 (curated, highest priority first)
122 words, every one with explicit evidence: a tested-question hit, high cross-administration frequency, or a question-stem function. The first rows are the most-tested missing targets; it also ports VocabLoco's one asset — its ~8 exclusive question-stem words (inference, summarize, contrast, comprehension…) — so those words carry over in a switch.
VocabLoco wastes 65% of slots (sight words, contractions, math jargon); Vocabulon's words are tested targets 2.4× as often
4
Vocabulon · strong
3–0
This page
5
Vocabulon · moderate
2–1
Owns the actual tested targets (isolated — the #1 g5 word, intricate, specimen); per-slot rate otherwise close
Appendix · Method & verification
How this was built and reviewed
Corpus2,191 released ELA documents · 25 state programs + 38 national/commercial assessments · grades 3–6 · 45M characters. Grade 4: 380 docs deduplicated to 70 unique administrations across 19 programs.
Tested-word mining978 explicit vocabulary-question targets extracted by pattern ("What is the meaning of ___ in paragraph 2?") — 167 at grade 4. Spot-checked sample: 25/25 genuine. Counts collapse shared item banks (IAR/PARCC, MCAS/RICAS, STAAR RLA/Reading), so a question duplicated across exports is counted once.
ClassificationAll 4,686 grade-4 corpus words + 634 list words rated for Beck tier and known-grade. Tested words dual-rated with tiebreaks; single ratings audited by 3 independent raters: 89% tier / 99% known-grade agreement.
MatchingStrict inflectional variants only (isolate ⇒ isolated) after the loose matcher was caught crediting quest for "question" — that artifact and all counts were corrected before judging.
JudgingThree independent lenses (coverage, difficulty/waste, path-to-guarantee) scored both lists blind to which is in-house. 3–0, scores 58–36, 72–38, 62–42.
Adversarial reviewTwo reviewers were instructed to overturn the verdict, one on statistics and one on decision logic. Both returned "not refuted" at high confidence, and the statistical recomputation widened the lead.
Caveats
As shipped, the better list covers 28.2% of tested grade-4 words. The guarantee comes from the add-list; the verdict is about which base to build on.
The grade-4-page-only tested slice is a tie (7.7% each). Vocabulon's headline edge comes partly from grade-3 placements, which the timing-corrected metric credits, and partly from grade-5 placements, which it excludes.
Vocabulon's kg-7+ tail (~46 words) is only partly justified. Its total waste is about 27%, against VocabLoco's 44% or more.
Classification carries an ~11% tier-disagreement band (audited), far too small to close a 34%-vs-5% junk gap that direct inspection confirmed.
Two inflation artifacts were caught in self-audit and corrected on this page: doc-level tested counts double-counted shared item banks (the "spectacular ×5" items are one PARCC/IAR question), and an early 3:1 timing-corrected ratio came from a subsample; the exact recomputation is 1.9× (20.5% vs 11.1%).
Released tests from 70 administrations predict future tests imperfectly. MAP releases no items, so the MAP claim rests on similarity to the state tests measured here.