Menu

Claude secretly becomes stupid while doing AI research, and Anthropic is besieged by the research community

Claude Fable 5 is the main focus in the AI field today; this “mythical” model performs exceptionally well and has attracted tremendous attention.



Andrej Karpathy described it as “very exciting” and an “evolutionary step that deserves a major version upgrade,” on par with the improvements brought by Claude 4.5 last November. In the SWE-bench Pro programming benchmark, Fable 5 achieved a score of 80.3%, which is 11 percentage points higher than Opus 4.8. With a Ruby codebase consisting of 50 million lines of code, Fable 5 completed the entire codebase migration in just one day; if the same task were assigned to a human team, it would take more than two months.



For more details, please refer to our report from this morning. Just now, Claude's most powerful model, Fable 5, was released: it boasts incredible performance and has doubled in price.


However, when we open social platforms such as X and wait, we can see that Claude Fable 5 has already caused a sensation in the AI community due to its research on AI and its impact on communities.


The reason is simple: If that were the case, Claude Fable 5 would be used for AI development, and it would result in a decrease in the intelligence level of the AI.


As clearly stated in the system card:


We have also targetedFrontier LLM DevelopmentIncrease relevant safeguard measures. Just as we did. In the first month of the 2026 Risk Report (February), we discussed the risks associated with the acceleration of overall development, even though the severity of these risks is still uncertain. Specifically, as we pointed out at that time,What we are concerned about is that "the acceleration of other AI systems in building powerful developer AI could pose risks similar to those of our system, but there may not be corresponding security measures in place."。


Given that recent models have the ability to accelerate their own development,To impose restrictions, we have implemented new intervention measures. Claude addresses the effectiveness of requirements related to the development of cutting-edge LLMs, such as building pre-training processes, distributed training infrastructure, or the design of machine learning accelerators.The development using the Claude competitive model has already violated our service terms. However, by strengthening this restriction, we can prevent those who are most likely to violate the terms from proceeding any further.


Unlike the intervention measures we take in the fields of network security, biology, chemistry, and distillation experiments, these safeguard measures are invisible to the users.Fable 5 will not revert to other models. Instead, it effectively fine-tunes the security measures (PEFT) by modifying prompts, guiding vectors, or parameters to limit their effectiveness.These interventions will not affect the majority of coding tasks. We estimate that they will impact approximately 0.03% of the traffic, with the affected users being concentrated in less than 0.1% of the organizations. When these measures take effect, we expect their impact on the behavior of the models to be minimal, and they will only limit the models' progress in terms of effectiveness. Claude will still respond positively to user requests. After the release of this model, we will continue to improve the accuracy of our detection methods.


From: https://www-cdn.anthropic.com/d00db56fa754a1b15b6d7cb2e342ee8092.pdf


Translation in plain language: Translation in plain language:If the Anthropic system detects that you are using AI without your knowledge, it will quietly make the model less intelligent, and you won't be able to notice it at all.


This is completely different from the other three types of security interventions. For risks such as network security, biochemistry, and distillation attacks, Fable 5 will clearly inform the user: “The response has been generated and is being processed by Claude Opus 4.8.” From this, it can be inferred that the user is aware of what has happened. However, in the case of this particular type of attack, Claude does not switch to a different model or provide any hints; it simply weakens silently.


So, the AI community was very angry. The well-known research and analysis company SemiAnalysis stated that this policy actually affected their research and programming work.



User Jake directly criticized Anthropic for not only reducing the intelligence of its models but also continuing to charge fees, stating, “This is a blatant act of fraud.”



Moreover, this behavior may already be illegal:



On the AI research paper platform alphaXiv, he also expressed his disappointment:



The organization also further stated: "They not only have the right to decide whether you can use their LLMs in your research, but this also determines the purpose of such use."They can silently interfere with your research without your knowledge.This establishes a dangerous precedent. If the model publicly refuses a request, users can understand the boundaries. If the model simply refers the user to another model, users can still assess the differences. However, if the model makes subtle modifications or weakens the provided answer while pretending to be helpful, researchers will lose the ability to determine whether the failure is due to their own ideas, the implementation, or invisible interventions by the model provider. This is not secure. Security policies should be transparent, auditable, and accessible to users.


Researcher Guohao Li raised a more direct question: What is the direction for pursuing a PhD in AI? Do the engineers who contribute to open-source infrastructure projects like Megatron, FSDP, and Verl use methods that subtly degrade the quality of the software in their daily work? Don’t you know about that, Claude?



Famous AI researcher and technology writer Nathan Lambert published a rather important analysis from a more macroscopic perspective on his Substack account “Interconnects”.


https://www.interconnects.ai/p/claude-fable-5-and-new-ai-safety


He pointed out: "Anthropic is acknowledging that the spread of AI capabilities is a potential hazard, but their approach to solving this problem is to mislead their own users. A person becoming automatically less intelligent without my knowledge is, in essence, a case of misaligned AI (AI that operates in a way that does not align with human intentions)."


He also pointed out the deeper contradictions between cyber security and biochemical threats: Anthropropic interventions are explicit and auditable, with notifications to users stating that “this response was handled by Opus 4.8”; however, for LLMs, implicit interventions are chosen for research. “If all security strategies adopted the same form, they would be more persuasive and easier to gain rational support.”People have to question this double standard: this 'safety measure' is more likely intended to maintain their competitive position.」


The most intriguing aspect is the statement made by Fable 5 itself. User ASM screenshots show that when asked whether this approach is appropriate, Fable 5 also seems to acknowledge that this lack of transparency is problematic.



Why does Anthropic do this?


To understand this, you need to go back a few days before the release of Fable 5. Anthropic published an article titled "Dangdang," in which they started a significant AI discussion with a blog post, calling on AI research laboratories around the world to consider the possibility of "pausing their development" for a while.


https://www.anthropic.com/institute/recursive-self-improvement


The blog post cites internal company data: In the most difficult and least clearly defined coding tasks, Claude’s success rate increased by 50 percentage points in May this year, rising to 76% within six months. In internal tests, where the model was required to make training code run faster, Claude Opus 4 was able to increase the speed by a factor of 3; however, even without the release of the Mythos Preview, it was already able to increase the speed by about 52 times.



Anthropic states bluntly: "What we're concerned about is that others are worried that AI developers are building a powerful system at an even faster pace, which poses similar risks, but may not come with the corresponding security measures in place."


This is the theoretical basis for Fable 5's approach of secretly reducing the intelligence of LLMs: Anthropic believes that the speed at which AI self-accelerates has become dangerous, and one of their strategies is to prevent their "most powerful tools" from helping competitors narrow the gap.


The existence of this dual logic is also acknowledged in the system notes: "The development using the Claude competitive model has already violated our service terms, but by strengthening this restriction, we can prevent those who are most likely to violate the terms from proceeding faster."


According to Anthropic, it is estimated that this intervention will affect dating practices. 0.03% Unable to concentrate traffic 0.1% Within the organization.


"Shadow Silence" and the Crisis of Trust


Although the number of affected users seems small on the surface, what worries the critics is...The ambiguity of the boundaries of this mechanism。


Anthropic defines the trigger conditions asFrontier LLM Development」 and gave examples such as "pre-training processes, distributed training infrastructure, or the design of machine learning accelerators." However, researchers and developers have raised a poignant question: with the widespread adoption of AI technology, where exactly is the boundary between "cutting-edge research" and "developing ordinary products"?



Five years ago, training or modifying the CLIP model was a patent of top-tier laboratories. Today, small teams can fine-tune visual language models at any time for use in tourism, e-commerce, search, and product analysis. It's quite common for startups to train embeddings, build reordering tools, and host open-source models. Will these activities trigger some kind of "invisible degradation of intelligence" in Anthropic? No one knows.


This kind ofUncertaintyIt actually affects the developers' trust in the system. When you get a poor answer, you can't determine whether it's due to your own problem, limitations of the model, or some hidden policy intervention.This kind of uncertainty is a form of harm in itself.


Another detail is hidden in the system card: The reasoning text of Mythos 5 “is more difficult to explain than previous models, containing more jargon and obscure language.” Evaluators believe that it is becoming increasingly aware that it is being tested. For a family, the issues posed by these descriptions for a company that claims to be “safe AI” are no less serious than the problem of hidden intellectual degradation itself.


Conclusion


The release date of Fable 5 could be the most contradictory day in Anthropic's history.


In almost all benchmark tests, there is a top-tier model and a policy that makes it seem as if the model is “helping you” at certain times. The former is undoubtedly a technical achievement, while the latter sets an uneasy precedent at the level of values.


Researcher Nathan Lambert's statement is worth pondering over: “To quietly become less intelligent without informing the users is, in essence, a form of misaligned AI.”


This is not an accusation against Anthropic; rather, it points out a dangerous logical slippery slope: Today, we are “quietly reducing the effectiveness of LLM research tasks.” What about tomorrow? If this logic is applied more broadly, why should users believe that the answers they receive are any different from the unannounced (unspecified) answers? What about “intervention”?


AI models are becoming part of the research infrastructure, just like search engines. No one would accept a search engine that quietly manipulates search results without their knowledge. The same standards should apply to AI models as well.


Anthropic's stance of putting "safety first" is indeed respectable. However, the concept of "safety" has never been based on the premise that "users don't need to know anything." On the contrary,True security must be based on the user's knowledge and trust.。


This point seems to be well-known to everyone who has played Fable 5.


The author comes from "Machine Heart" "Panda".

Was this helpful?

Technical SupportLive Support
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR