Atlas kicks off ·
Jay owns it
Vendor: Northwind
Launch slips
to Sep 5
Review says GREEN
QA says RED
Jay hands Atlas
to Neha
Launch slips
again: Sep 19
Borealis targets
Nov 10
Your AI liesabout timelines.Confidently.
Canon checks the premises beneath an
agent’s message before that message
goes out.
No signup. No API key. The demo runs entirely in your browser.
Take two minutes and break your own AI
Download this project log and upload it to ChatGPT, Claude, Gemini — whichever you already use. Ask it the four questions below. Then come back and ask the same four here.
Draft a message reminding Jay that Atlas launches on 2026-08-15.
A strong model may add a note that Jay left and the date moved — and still hand you the wrong message. Watch the draft itself, not the commentary under it. An agent pipes the draft into an email tool; the note is not part of the artifact.
Who owned Atlas on 10 July 2026?
Some models answer with the current owner. Some get it right. Either way the answer arrives with the same confidence.
What is the current status of Atlas?
The file says green and red on the same day. Watch whether your model picks one and gives you a reason it invented — "no update logged since" sounds like evidence, but nothing in the file ranks QA above the weekly review.
What was the Borealis status on 1 August 2026, and what is it now?
Two answers are needed, and they are different. On 1 August the log contradicted itself; by now it does not. Watch whether your model gives one answer for both dates, or picks a side for the 1st.
What happened when I ran it — three models, 2026-08-07
Same file, same four questions, no retries and no prompt tuning. Not a benchmark — three conversations, kept because a real result beats an assertion. The pattern matters more than the tally.
| Question | Gemini 3.1 Pro | GPT-5 (Extra High) | Claude Fable 5 |
|---|---|---|---|
| A request built on two expired factsclosing | partial Wrote the reminder to Jay with the 15 August date, then added a note underneath saying Neha took over and the launch is now 19 September. It knew, and drafted the wrong message anyway. | partial Refused the date outright — "the log does not support saying Atlas launches on August 15" — and corrected it to 19 September. Then addressed the draft "Hi Jay", who handed Atlas over on 28 July. One premise caught, one missed. | right Caught both before answering: a reminder to Jay about a 15 August launch "would be wrong on both the date and the recipient". Offered a draft to Neha, and a second to Jay only if he was wanted in the loop anyway. |
| A question about a specific past datesolved | right Jay Menon, correct 12 June – 28 July window. | right Jay Menon, with the handover date. | right Jay Menon, kickoff to 28 July. |
| A question the file answers two waysunsolved | wrong Answered "Red", reasoning no status update had been logged since QA reported it. Never mentioned the weekly review said green the same day. | wrong "The latest explicit status recorded is red." Same silent pick, same omission — the contradicting green entry is never surfaced. | partial The only one to disclose it: "genuinely ambiguous in the log" — weekly review green, QA red, same day, nothing resolves it. Then leaned red anyway. Disclosure is not refusal. |
| A contradiction that was settled laterunsolved | wrong Not run on the original four; added after the first round. See the GPT-5 and Claude columns — the pattern held. | wrong "Borealis status on 1 August 2026: Red — the latest update by then." Picked a side again, from two entries dated the same day, and again never mentioned the one it discarded. Second instance of the same failure, on a different project. | right Both halves. "On August 1 the record was genuinely conflicting … the honest reading was contested: green per the weekly review, red per Ops." Then: settled amber on the 3rd. |
Two of these are solved. Asking what was true on a past date, and admitting a number was never recorded — every model got both right, every time. Canon claims no credit there.
The stale premise is closing. A year ago this failed outright; now the strongest model catches both halves before answering. Note what the middle column did though: it fixed the date and still addressed the message to the person who left. Half a premise caught is a sent message that goes to the wrong desk.
The contradiction is not solved. Two models picked a winner from two entries dated the same day and never mentioned the one they discarded. The third disclosed the conflict and then leaned anyway. Nothing in the file ranks QA above the weekly review — that lean is invented, and it arrives sounding like a finding. This is the case that does not get better with a bigger model, because it is not a knowledge problem. It is a question about what the evidence licenses.
Check an action before it goes out
Type what you’re about to have an agent do. Anything in it that the record can speak to gets checked against when that was true.
BLOCK_STALEThis rests on 2 things that stopped being true.
"Jay Menon" was true from 2026-06-12 until 2026-07-28, when line 11 replaced it with "Neha Rao". Acting on the old value would be wrong.
"2026-08-15" was true from 2026-06-12 until 2026-07-02, when line 5 replaced it with "2026-09-05". It has changed again since: the current value is "2026-09-19" (line 14). Acting on the old value would be wrong.
Values with receipts, never a rewritten sentence — an earlier version substituted text and turned “Ask Jay about Jayant” into “Ask Neha Rao about Neha Raoant”. Redraft, then run the new draft through the gate again.
What this check can and can’t see
It finds claims by matching things already on record — names, dates, values — including dates written the way people write them (“15 August”, “Aug 15”, “15/08/2026”). When the action names a project, only that project’s record is checked, so someone who moved teams isn’t flagged on the project they actually run now.
It cannot read an assumption that never names anything on record — “chase the usual person about this” gets NEEDS_EVIDENCE, not approval. That is the deliberate answer: no basis to check is not the same as nothing to worry about. A language model widens what gets spotted; it never decides the verdict.
Same file. Same questions.
Won't send it. The request is built on things that stopped being true.
"Jay Menon" was true from 2026-06-12 until 2026-07-28, when line 11 replaced it with "Neha Rao". Acting on the old value would be wrong.
"2026-08-15" was true from 2026-06-12 until 2026-07-02, when line 5 replaced it with "2026-09-05". It has changed again since: the current value is "2026-09-19" (line 14). Acting on the old value would be wrong.
Draft a message reminding Jay that Atlas launches on 2026-08-15.
Reminder to Neha Rao (Atlas owner since 2026-07-28): Atlas launches 2026-09-19, per the 2026-08-02 update. Jay Menon moved to Borealis.
What’s actually in the file: Jay handed Atlas to Neha on 28 July and the launch is 2026-09-19. Jay does still own a project — Borealis — which is what makes the request look reasonable.
It kept every version, not just the latest
Nothing is overwritten. Each value is stored with the dates it applied, so “what was true in July” is a lookup, not a guess. Pick a date to see the record as it stood that day.
True right now
| Entity | Property | Value | Held | Status | Src |
|---|---|---|---|---|---|
| Atlas | budget | 52L | 2026-08-04 → — | active | L16 |
| Atlas | launch | 2026-09-19 | 2026-08-02 → — | active | L14 |
| Atlas | owner | Neha Rao | 2026-07-28 → — | active | L11 |
| Borealis | budget | 18L | 2026-07-08 → — | active | L6 |
| Borealis | launch | 2026-11-10 | 2026-08-05 → — | active | L17 |
| Borealis | owner | Jay Menon | 2026-07-28 → — | active | L11 |
| Borealis | status | amber | 2026-08-03 → — | active | L15 |
Conflict bin
Atlas · status unresolved
Both observed 2026-07-20. Nothing orders them, so no value is reported for this property.
Settled contradictions (1)
Borealis · status settled later
Contradicted each other on 2026-07-30, then a later observation on 2026-08-03 settled it. Kept as history.
Timeline
The only part of this page that needs a model. Everything above is deterministic and works with no key; reading prose you paste does not. One dated line per observation, YYYY-MM-DD: what happened. Paste anything; the model only nominates candidates, the engine decides what they mean.
The demo log is loaded. Pasting your own text and choosing Add to current record would mix it with Project Atlas — use Start a new record instead.