the word-list case · grade 4 · g5 · g6

The Grade-4 Case

Grade 4 is the deepest slice of the corpus: 380 documents, 70 unique test administrations, 19 programs. This brief measures both lists against all of it.

"The right word list is the correct word list that allows them to deliver a 90%+ on any state or Common Core reading comprehension test… You just pull down tests and get all the state tests and see if your word list is going to teach all of those." — Andy, Aug 27 demo (the approval criterion)
Verdict: Vocabulon, by a strong margin.
Unanimous 3–0 judge panel across coverage, difficulty, and guarantee lenses. The margin held through adversarial review. A second audit caught two inflation artifacts, item-bank duplication and a subsample ratio, and both corrections appear below.
3–0
lens panel, all three for Vocabulon (2 strong, 1 moderate)
1.9×
lead on tested words taught by test day (20.5% vs 11.1%)
0/2
adversarial refutations landed; both upheld the verdict at high confidence
4,686
grade-4 corpus words tier-classified to find every tier-2/3 target
Before the exhibits · The naive count

The quick spreadsheet check misleads in both directions

Counting raw word overlap produces three plausible-looking numbers. All three ignore the same question: does the word teach a grade-4 kid anything before test day?

Three naive counts and what's wrong with each
L = VocabLoco · U = Vocabulon
Naive metricVocabLocoVocabulonWhy it misleads
Tested words on each product's grade-4 page only7.7%7.7%A dead tie, because it ignores grade placement: grade-3 words are already taught by test day, grade-5 words are not
STAAR G4 explicit vocab targets covered (n=25)44Tie again, and 18 of 25 are on neither list. The STAAR story is the gap
List words found anywhere in test docs, frequency-weighted9.96.3VocabLoco "wins" here because because, although, and internet match text everywhere while teaching nothing. The slice is also inflated by test boilerplate

Even so, the simplest defensible count (raw any-grade overlap with tested words, no filters at all) still picks Vocabulon, 27.2% vs 16.8%. No counting method we could construct flips the winner. The exhibits below show why.

Exhibit A · What tests ask

Coverage of the words tests explicitly test

We regex-mined every "What is the meaning of ___?" style question across all 2,191 documents: 167 grade-4 vocabulary targets, of which 117 are teachable tier-2/3 words (words a majority of grade-4 students don't already know). Explicit targets are the strongest signal: these are the words test writers choose.

VocabulonVocabLoco
Tested tier-2/3 grade-4 words covered by each list
percent of tested words on the list · strict inflectional matching (isolate ⇒ isolated)
On the list at any grade 3–5 · n=117 teachable targets
Vocabulon
28.2%
VocabLoco
16.2%
Taught by test day: only grade ≤ 4 placements count, which is how school sequences the words
Vocabulon
20.5%
VocabLoco
11.1%
data table
Metricn / weightVocabLocoVocabulon
Any grade 3–5, all tier-2/3 tested (raw)12516.8%27.2%
Any grade 3–5, teachable (kg≥4)11716.2%28.2%
Timing-corrected (grade ≤4 placements)11711.1%20.5%
Words covered (of 117 teachable)1171933

The timing correction was the adversarial reviewers' strongest objection to the headline number: "any grade 3–5" credits grade-5 placements a grade-4 tester has never been taught. Recomputed exactly, Vocabulon holds 20.5% vs VocabLoco's 11.1%, still nearly 2:1, because six of VocabLoco's nineteen hits (abruptly, retreat, transformation…) sit only on its grade-5 page.

The 30 most-tested grade-4 words, and who has them

spectacularNEITHER dinNEITHER shrieksNEITHER abruptlyL·g5 anonymousNEITHER applaudedNEITHER chaosU·g4 composureNEITHER deckedNEITHER dominanceNEITHER dwindlesNEITHER encumberedNEITHER idleNEITHER insulationNEITHER obligeNEITHER precedesNEITHER retreatL·g5 solitaryNEITHER transformationBOTH·g5 adaptedU·g3 affixNEITHER aggressiveU·g5 applaudNEITHER channelNEITHER concealedNEITHER disputingU·g5 enforceNEITHER extraordinaryNEITHER foragingNEITHER humbleL·g3

Ordered by how widely each word was tested. After collapsing shared item banks (IAR/PARCC/NJSLA, MCAS/RICAS, STAAR variants), nearly every word here is one unique test item, so the unit of counting is the word rather than the document. L/U shows which list carries the word and at what grade. The words with no tag on either list are the subject of Exhibit D.

Exhibit B · Where the slots go

How each list spends its teaching slots

Every one of the 634 grade-4 list words was classified by Beck tier and by the grade at which a majority of US students already know it (dual-rated for tested words, independently audited at 89% tier / 99% known-grade agreement). Here is what each list spends its slots on:

VocabLoco · 368 grade-4 words194 productive slots
34% junk
10%
53% sweet spot
Vocabulon · 266 grade-4 words195 productive slots
5%
73% sweet spot
17%
junk (tier 0/1) already known by g3 sweet spot (kg 4–6) stretch (kg 7+)

The two lists carry an almost identical number of productive words, 194 and 195. VocabLoco needs 102 extra slots of a child's practice time to deliver the same payload. 44% of its grade-4 list teaches nothing a typical fourth-grader doesn't already have.

What's on VocabLoco's grade-4 list

A sample of the 125 tier-0/1 entries: sight words, function words, and spelling-drill items no ELA test asks the meaning of.

becausealthoughalreadybuildcompletedifferentexplainfactfinallyimportantinternetislandkeyacceptadvicebalancebelievecontinueeasilyenergyfamoushappilydesert / dessertcalculatorautomobilecostumebiscuiteighth
Vocabulon's own weak tail: 46 words (17.3%) sit at known-grade 7+, mostly social-studies curriculum vocabulary (feudalism, shogun, serf, ratify, dynasty, monastery). Some of that is justified stretch, since tests do target kg-6–8 words, and part is waste for ELA test prep. Total waste comes to roughly 27%, against VocabLoco's 44% or more.
Exhibit C · The passages

Passage vocabulary: a smaller edge

Beyond explicit vocab questions, kids meet hard words inside reading passages. Among the 900 teachable tier-2/3 words appearing in 3+ of the 70 unique grade-4 test administrations, Vocabulon covers 24.2% vs VocabLoco's 18.1%. The "convergence method" (words appearing across N independent test programs — the method already used for the grade-3 ITBS analysis) gives the same picture at every threshold:

Convergence coverage — tier-2/3 words appearing in ≥ K of 19 grade-4 test programs
≥ K programswordsVocabLocoVocabulon
32,12512.2%13.9%
41,28113.3%15.1%
580314.7%16.4%
650517.4%18.4%
819620.4%19.9%
Near parity, with a slight Vocabulon edge. Grade 4 is not the grade-3 blowout (74% vs 34% on ITBS), which is why the slot-quality and tested-target evidence above carries the verdict.
Exhibit D · The gap neither list covers

Neither list gets near 90% today

67%

of the 117 teachable explicitly-tested grade-4 tier-2/3 words — including (spectacular, din, shrieks, applauded, composure) — are on neither list.

The "guarantee we cover everything previously tested" therefore has to be delivered by adding words. That turns the decision question into: which list is the better base to patch?

Vocabulon as chassis
  • Starts at 28.2% tested-word coverage
  • 270 curated adds close the gap
  • Nothing to remove (5% junk)
  • Selection already tracks what tests target
VocabLoco as chassis
  • Starts at 16.2% tested-word coverage
  • 324 adds needed, a bigger patch
  • Plus ~a third of the list should be pruned
  • Ends up rebuilt into a different product

The recommended add-list for Vocabulon grade 4 (curated, highest priority first)

122 words, every one with explicit evidence: a tested-question hit, high cross-administration frequency, or a question-stem function. The first rows are the most-tested missing targets; it also ports VocabLoco's one asset — its ~8 exclusive question-stem words (inference, summarize, contrast, comprehension…) — so those words carry over in a switch.

spectaculardinshrieksabruptlyidleretreatsolitaryanonymouscomposureapplaudconcealedextraordinaryinferencesummarizecontrastphrasequotationexcerptsternmerelyintendlimbsidiomoriginallysecureinsulationdominancedwindlesprecedesmimicmetamorphosisenforceforagingwitherstabilizeinquired
all 122 recommended additions

spectacular · din · shrieks · abruptly · idle · retreat · solitary · limbs · idiom · intend · merely · originally · secure · stern · responses · document · punctuation · phrase · select · quotation · inference · excerpt · refer · accurately · contrast · quotations · purposes · statements · comparison · trait · encourage · deemed · efficiently · frequency · summarize · comprehension · anonymous · composure · dominance · dwindles · precedes · insulation · impressed · extraordinary · humble · inquired · stabilize · wither · applaud · concealed · enforce · foraging · metamorphosis · mimic · pleading · pace · snarled · conserve · exhausted · enthusiastic · sincere · summit · hoisted · invincible · jolted · lunged · misfortunes · mountainous · patrol · pure · ration · recollection · refining · submerged · grammar · essay · capitalization · identical · communicate · demonstrate · displayed · prompt · regardless · textual · aspects · expectations · graphic · demonstrates · indicate · selections · unfamiliar · contrasting · experiences · relate · undergo · decked · oblige · affix · stowed · irritating · perplexity · gnawed · hurled · annual · flaw · muffles · naysayers · procure · retraces · designated · vocabulary · comprehend · portions · displays · organizational · applying · exception · involves · obtained · transition · insight · judgment

Context · The other grades

Grade 4 next to the other grades

GradeVerdictPanelOne-line reason
3Vocabulon · strong3–0VocabLoco wastes 65% of slots (sight words, contractions, math jargon); Vocabulon's words are tested targets 2.4× as often
4Vocabulon · strong3–0This page
5Vocabulon · moderate2–1Owns the actual tested targets (isolated — the #1 g5 word, intricate, specimen); per-slot rate otherwise close
Appendix · Method & verification

How this was built and reviewed

Corpus2,191 released ELA documents · 25 state programs + 38 national/commercial assessments · grades 3–6 · 45M characters. Grade 4: 380 docs deduplicated to 70 unique administrations across 19 programs.
Tested-word mining978 explicit vocabulary-question targets extracted by pattern ("What is the meaning of ___ in paragraph 2?") — 167 at grade 4. Spot-checked sample: 25/25 genuine. Counts collapse shared item banks (IAR/PARCC, MCAS/RICAS, STAAR RLA/Reading), so a question duplicated across exports is counted once.
ClassificationAll 4,686 grade-4 corpus words + 634 list words rated for Beck tier and known-grade. Tested words dual-rated with tiebreaks; single ratings audited by 3 independent raters: 89% tier / 99% known-grade agreement.
MatchingStrict inflectional variants only (isolate ⇒ isolated) after the loose matcher was caught crediting quest for "question" — that artifact and all counts were corrected before judging.
JudgingThree independent lenses (coverage, difficulty/waste, path-to-guarantee) scored both lists blind to which is in-house. 3–0, scores 58–36, 72–38, 62–42.
Adversarial reviewTwo reviewers were instructed to overturn the verdict, one on statistics and one on decision logic. Both returned "not refuted" at high confidence, and the statistical recomputation widened the lead.

Caveats