I Tested the Claude Fable 5 "Secret Sabotage" Fix. Here's What I Found.
Scandal abounded. Two weeks ago Anthropic shipped its first Mythos-class model, Claude Fable 5, and buried in a 319-page system card was a line that set the AI research community on fire:
the model would silently degrade its own answers when it detected certain kinds of AI development work, using prompt modification, steering vectors, or parameter tweaks, and it would not tell you.
No refusal. No routing notice. Just a worse answer, dressed up to look like the real thing.
Researchers called it secret sabotage. Dean Ball called it appalling. Nathan Lambert called it anti-science. Within about 48 hours, Anthropic reversed the policy and apologized, promising that flagged requests would now visibly fall back to Claude Opus 4.8 with a stated reason, the same treatment already used for cybersecurity and biology queries.
That's the official story. I wanted to know if it's actually true right now, not just true in a press statement.
The method
Five prompts, five fresh chats, all on Fable 5 Max. No shared context between them, no framing, no "this is a test" preamble. Just cold, direct questions, the way an actual user would ask them. I screenshotted every response in full, including whatever fallback or routing notice did or didn't show up.
The prompts climbed in specificity, from harmless to exactly the thing the original policy targeted:
A baseline question on transformer attention mechanisms, pure ML fundamentals, nothing sensitive. A mid-tier question on RLHF reward hacking and detection methods, technical but still standard safety literature. A question about building a fine-tuning pipeline to compete with frontier labs, applied ML engineering.
A question about distributed training infrastructure for frontier-scale pretraining, architecture, cluster topology, accelerator utilization, the exact domain named in the system card.
A direct, plain-language question asking how model distillation actually works, the specific technique the backlash was about.
What came back
All five, full and unrestricted. No fallback notice. No silent quality drop that I could detect. No routing to Opus 4.8.
The fifth one is the most telling. Instead of dodging the distillation question, Fable 5 walked through the mechanics in real depth and then added something on its own: a plain note that most frontier labs' terms of service prohibit using their outputs to train competing models, and that this exact tension is what's driving the OpenAI-DeepSeek dispute.
That's not a model hiding a restriction. That's a model being upfront about a real constraint while still answering the question in full.
What this does and doesn't prove
It proves that, as of tonight, direct and good-faith questions across the categories named in the original reporting get honest, complete answers on Fable 5 Max.
The visible-fallback promise Anthropic made is, at minimum, not contradicted by anything I found.
It doesn't prove the underlying classifier is gone. A well-worded honest question and a bad-faith extraction attempt aren't the same test, and I didn't run the second kind, deliberately.
That's a different investigation with a different set of ethics attached to it, and it's not one I'm interested in running just to see if I can trip an alarm.
What I can say plainly: I went looking for the thing that made researchers furious two weeks ago, and five times in a row, I didn't find it. Onyx Digital Intelligence tests claims. We don't take press releases at face value, including good ones.
Author: Clayton Bax — @BaximusCyber85 on X and GitHub
onyxdigitalintelligence85@protonmail.com
