408 items that passed every automatic check and both content gates: one of each kind of material for each grade band and course tested, across 24 courses from Kindergarten to Grade 12. Below them, the protections every item passes through, and the evaluation runs that measured them.
These are the strongest examples from the evaluation runs, chosen by the checks, not by hand. Verified means every automatic check and both content gates passed; Reviewed means a signed-in reviewer approved it as well. Open any item to see its checks, its sources and its downloads. The full gallery holds every example, including the ones held for review.
Each cell is the number of examples, and under it how many pass every automatic check and content gate. 0 of 408 items are held for review. Click a cell to filter the gallery.
Curriculum only. Every item starts from retrieved curriculum text (the outcomes a person picked, or the staged selection for a plain-language request). The writer sees only those sources and cites them by alias; no web content, no memory of other lessons.
Staged selection. A plain request such as "Grade 3 math and English for my son" is turned into intent, scope, sets and an outline before anything is written, and the person confirms the outcomes.
Archived programs excluded. Withdrawn curricula stay out of search, selection and generation unless asked for, and are always labelled.
Grade-band register. Vocabulary, sentence length and picture rules follow the audience band (K to 3, 4 to 6, 7 to 9, 10 to 12). The grade the person asks for decides the band; the highest grade among the sources is only the fallback. Senior pictures are diagrams and equipment with no people.
While writing
Main tier where facts are fragile. Detailed packs, presentations and family overviews are written on the main model tier at least; the fast tier held four times as many family overviews for wrong course facts.
Structured output. Each kind of material is written to a schema and validated; a malformed answer is repaired once from the model, never patched by hand.
Deterministic puzzles and diagrams. Word searches, crosswords, matching games, hangman and mazes are built by code from the extracted vocabulary and checked as solvable, and their rules, word lists and answer keys are written by code from the board that was built; scientific diagrams are drawn by code from a validated specification.
Solver-verified brain teasers. A logic puzzle is built by a constraint solver from a scenario the writer supplies: code picks the solution, writes clues that admit exactly that solution, orders the hints and prints the key with the reasoning. Categories linked in fact (a solid, vibrating particles, the strongest forces) keep their true pairings. The clues, hints, key and rules belong to the solver and are written back after every text repair, so no rewrite can turn a clue false.
Slide citations as pills. A slide bullet ends with its citations as the same pills every other kind shows; duplicates and empty brackets are removed at generation, and the PDF prints the pills in the slide footer.
Plain mathematics. Formulas, units and scientific notation are written as plain text in the sentence, never as code or LaTeX.
Pictures checked by a vision model. Every generated picture is read back by a vision model against the scene, the topic, the no-text rule and the audience rule; a rejected picture is redrawn once and a picture still rejected is never served. A safety-filter block is redrawn as a no-people diagram.
Citations resolved server-side. The writer uses short aliases; the server turns them into links to the exact outcome, so a citation can never point at something that does not exist.
After writing: 32 automatic checks on every item
Check
Severity
Applies to
What it checks
schema_valid
blocking
every kind
The stored body and parameters match the kind's schema.
question_count_matches
major
Quiz, Homework sheet
A quiz or sheet has exactly the number of questions requested.
every_question_has_answer
blocking
Quiz, Homework sheet
Every question has an entry in the answer key.
item_types_match
minor
Quiz
Every requested item type (multiple choice, true or false, numeric, short answer) appears with its count.
slide_count_and_layouts
major
Presentation
A presentation has the requested slides, each with a known layout.
slides_plain_text
major
Presentation
No slide text carries citation residue, a raw alias or a doubled pill. The repair tidies it with no model call.
game_word_list_present
blocking
Game
A puzzle carries its word list.
alt_text_on_images
major
every kind
Every picture has alternative text.
no_latex_or_fences
blocking
every kind
No LaTeX, code fences or HTML leak into the text.
no_aliases_or_markers
blocking
every kind
No raw citation aliases or image markers remain in the body.
no_em_dashes
minor
every kind
House style: no long dashes.
no_self_talk
blocking
every kind
No model self-talk ("As an AI", "Let me", "I will now").
no_internal_codes_in_family_output
major
every kind
Family material carries no outcome codes or internal labels.
language_matches_request
blocking
every kind
The body is in the language requested.
senior_wording
major
every kind
Senior material never addresses "kids" or "boys and girls".
readability_in_band
major
every kind
Student-facing prose reads at the band's grade level (Flesch-Kincaid); teacher documents have an adult ceiling; a read-aloud story may sit two grades above the band, since listening runs ahead of reading.
citations_resolve
blocking
every kind
Every citation points at a curriculum node that exists.
citations_in_scope
major
every kind
Every citation stays inside the courses the request was about.
citations_match_text
major
every kind
A sentence that paraphrases one source cites that source, not a neighbouring one; the repair re-points it.
overview_rows_in_selection
major
Parent overview
A year-at-a-glance table lists only outcomes from the selected course.
selected_outcomes_referenced
major
Lesson, Lesson package, Presentation
A lesson references every outcome the person selected; a whole unit picked at once (more than eight) passes at four in five, and a deck is asked for about one outcome per content slide, with the ones left out named.
Every arithmetic statement in an answer key is recomputed and must be right.
puzzle_solvable
blocking
Game
Every puzzle is solved by the validator before it is served.
diagram_spec_valid
blocking
every kind
A code-drawn diagram's specification is complete and consistent.
wordmap_legend_consistent
major
Word map
A word map's legend matches its groups and colours.
image_verdicts_accepted
major
every kind
Every served picture passed the vision check.
no_people_in_senior_images
blocking
every kind
Senior-band pictures show no people.
png_for_every_svg
major
every kind
Every SVG board has a PNG twin for download and printing.
audio_duration_proportional
major
every kind
A read-aloud story's narration length matches its script.
exports_render
blocking
every kind
Markdown, DOCX and PDF exports render without error.
claims_supported
blocking
every kind
Content gate: every factual sentence is checked against the curriculum passage it cites (or, for a course-level item, the outcome beneath the cited node that it matches best). Any unsupported sentence fails.
judge_review
blocking
every kind
Content gate: an independent reviewer on a different model reads the whole item against the request and the passages: factual errors, grade fit, answer key and procedures, safety and bias. Any factual error, a wrong key or procedure, grade fit under 3 or a safety issue fails.
Repairs
Unsupported sentences. rewritten from their passages, then re-checked
Reviewer corrections. applied sentence by sentence, then a second review; answer-key fixes go to the main tier
Student prose too hard. simplified for the band without changing numbers; citations are swapped for tokens during the rewrite and restored afterwards, so no link is altered or lost
Missing quiz item types. added or replaced to match the requested mix
A rejected picture. redrawn once through the same pipeline and vision check
A wrong citation. re-pointed at the source the sentence matches, with no model call
Long sentences. split at their own clause boundaries first (a semicolon, a comma before and, but, so), with no word changed; the model is asked only when that is not enough
Long dashes and missing picture descriptions. replaced with commas and periods, and filled from the scene the picture was drawn from, with no model call
A picture rejected twice. withdrawn: the material goes out without it rather than with a wrong picture or held for one
A deck with the wrong number of slides. adjusted to the count requested; existing slides stay as written
Selected outcomes never reached. a short cited section connects the material to each of them
Puzzle sections after any repair. the clues, hints, key and rules are written back from the solver and the board, so a text repair can improve the introduction but never touch the puzzle
Up to five revisions. a failing check with a repair is repaired and re-checked until it passes or a revision changes nothing; the person sees each step
One fresh draft. a blocking failure that survives its revisions is written again from the same request and goes through its own revisions
Anything that survives. held for review, and queued for a fresh attempt in the background: a recovery worker regenerates the item with the same request and the full loop, twice at most, spaced out; a Verified result replaces the draft and the page says so. The card names what the checks found meanwhile.
Trust levels
Draft. unedited AI output, or something failed; read it before using it
Verified. every automatic check and both content gates passed; no person has read it
Reviewed. a signed-in teacher approved it after every gate was green; only Reviewed material can be made public by its owner
Evaluation runs
Each run generates a matrix of items through the same API the Create page uses, with the content gates on, and scores every item. "Pass all" means every blocking and major check and both gates passed on the first attempt, before any repair or review.
Finished
Run
Items
Pass all
Held
Claims
Reviewer
Failed to generate
2026-09-17 08:24
2026-09-17T06-50-58-125Z
218
60.1%
74
83.5%
78.4%
2
2026-09-17 07:50
2026-09-17T06-31-43-509Z
179
47.5%
72
80.4%
74.3%
1
2026-09-17 06:35
2026-09-17T05-41-23-921Z
156
71.2%
35
84.2%
85.5%
0
2026-09-17 04:32
2026-09-17T03-30-21-135Z
152
57.2%
62
74.3%
79.6%
0
2026-09-17 02:37
2026-09-17T02-13-30-578Z
156
75.6%
1
·
·
0
A
Welcome to the AILI Prototype
This application is an early-stage prototype being evaluated by a select group of education stakeholders. The purpose of this review is to gather feedback that will help inform future development and improvements.
Please note:
This prototype is intended only for individuals who received an invitation to participate in the evaluation and feedback process from the Government of Alberta.
Features, functionality, content, and responses may change and may not reflect a final production version.
Information generated by the prototype should be reviewed carefully and does not represent official government policy or advice.
The Government of Alberta does not endorse, condone, or take responsibility for any outputs generated through deliberate misuse or manipulation of this tool.
Your use of the prototype and any feedback you provide may be used to support ongoing evaluation and improvement efforts.
By selecting Continue, you acknowledge that:
You are participating in the evaluation of this prototype, or have received access through an official invitation to participate.
You understand that this is a test environment and that the prototype may contain errors, omissions, or incomplete functionality.
You will provide feedback through the designated survey process.