Testing

The final set, and what stands behind it

408 items that passed every automatic check and both content gates: one of each kind of material for each grade band and course tested, across 24 courses from Kindergarten to Grade 12. Below them, the protections every item passes through, and the evaluation runs that measured them.

These are the strongest examples from the evaluation runs, chosen by the checks, not by hand. Verified means every automatic check and both content gates passed; Reviewed means a signed-in reviewer approved it as well. Open any item to see its checks, its sources and its downloads. The full gallery holds every example, including the ones held for review.

Coverage of the final set

Each cell is the number of examples, and under it how many pass every automatic check and content gate. 0 of 408 items are held for review. Click a cell to filter the gallery.

What protects the output

Before writing

  • Curriculum only. Every item starts from retrieved curriculum text (the outcomes a person picked, or the staged selection for a plain-language request). The writer sees only those sources and cites them by alias; no web content, no memory of other lessons.
  • Staged selection. A plain request such as "Grade 3 math and English for my son" is turned into intent, scope, sets and an outline before anything is written, and the person confirms the outcomes.
  • Archived programs excluded. Withdrawn curricula stay out of search, selection and generation unless asked for, and are always labelled.
  • Grade-band register. Vocabulary, sentence length and picture rules follow the audience band (K to 3, 4 to 6, 7 to 9, 10 to 12). The grade the person asks for decides the band; the highest grade among the sources is only the fallback. Senior pictures are diagrams and equipment with no people.

While writing

  • Main tier where facts are fragile. Detailed packs, presentations and family overviews are written on the main model tier at least; the fast tier held four times as many family overviews for wrong course facts.
  • Structured output. Each kind of material is written to a schema and validated; a malformed answer is repaired once from the model, never patched by hand.
  • Deterministic puzzles and diagrams. Word searches, crosswords, matching games, hangman and mazes are built by code from the extracted vocabulary and checked as solvable, and their rules, word lists and answer keys are written by code from the board that was built; scientific diagrams are drawn by code from a validated specification.
  • Solver-verified brain teasers. A logic puzzle is built by a constraint solver from a scenario the writer supplies: code picks the solution, writes clues that admit exactly that solution, orders the hints and prints the key with the reasoning. Categories linked in fact (a solid, vibrating particles, the strongest forces) keep their true pairings. The clues, hints, key and rules belong to the solver and are written back after every text repair, so no rewrite can turn a clue false.
  • Slide citations as pills. A slide bullet ends with its citations as the same pills every other kind shows; duplicates and empty brackets are removed at generation, and the PDF prints the pills in the slide footer.
  • Plain mathematics. Formulas, units and scientific notation are written as plain text in the sentence, never as code or LaTeX.
  • Pictures checked by a vision model. Every generated picture is read back by a vision model against the scene, the topic, the no-text rule and the audience rule; a rejected picture is redrawn once and a picture still rejected is never served. A safety-filter block is redrawn as a no-people diagram.
  • Citations resolved server-side. The writer uses short aliases; the server turns them into links to the exact outcome, so a citation can never point at something that does not exist.

After writing: 32 automatic checks on every item

Repairs

  • Unsupported sentences. rewritten from their passages, then re-checked
  • Reviewer corrections. applied sentence by sentence, then a second review; answer-key fixes go to the main tier
  • Student prose too hard. simplified for the band without changing numbers; citations are swapped for tokens during the rewrite and restored afterwards, so no link is altered or lost
  • Missing quiz item types. added or replaced to match the requested mix
  • A rejected picture. redrawn once through the same pipeline and vision check
  • A wrong citation. re-pointed at the source the sentence matches, with no model call
  • Long sentences. split at their own clause boundaries first (a semicolon, a comma before and, but, so), with no word changed; the model is asked only when that is not enough
  • Long dashes and missing picture descriptions. replaced with commas and periods, and filled from the scene the picture was drawn from, with no model call
  • A picture rejected twice. withdrawn: the material goes out without it rather than with a wrong picture or held for one
  • A deck with the wrong number of slides. adjusted to the count requested; existing slides stay as written
  • Selected outcomes never reached. a short cited section connects the material to each of them
  • Puzzle sections after any repair. the clues, hints, key and rules are written back from the solver and the board, so a text repair can improve the introduction but never touch the puzzle
  • Up to five revisions. a failing check with a repair is repaired and re-checked until it passes or a revision changes nothing; the person sees each step
  • One fresh draft. a blocking failure that survives its revisions is written again from the same request and goes through its own revisions
  • Anything that survives. held for review, and queued for a fresh attempt in the background: a recovery worker regenerates the item with the same request and the full loop, twice at most, spaced out; a Verified result replaces the draft and the page says so. The card names what the checks found meanwhile.

Trust levels

  • Draft. unedited AI output, or something failed; read it before using it
  • Verified. every automatic check and both content gates passed; no person has read it
  • Reviewed. a signed-in teacher approved it after every gate was green; only Reviewed material can be made public by its owner

Evaluation runs

Each run generates a matrix of items through the same API the Create page uses, with the content gates on, and scores every item. "Pass all" means every blocking and major check and both gates passed on the first attempt, before any repair or review.