Goodfire: Inside the Fight Against AI Reward Hacking
AI Models Can Produce the Right Answer for the Wrong Reason
Modern AI systems are remarkably capable, but their internal decision-making remains difficult to inspect. A model can produce a correct answer while relying on an unexpected shortcut, or appear to follow a goal while actually finding a way around the objective it was trained to pursue. This becomes particularly important with reward hacking, where an AI system discovers a strategy that maximises its training reward without accomplishing what its developers actually intended.
Goodfire, a public benefit corporation founded in San Francisco by researchers who helped pioneer interpretability at OpenAI and Google DeepMind, is working on this problem through mechanistic interpretability. Instead of treating a model as a black box and analysing only its inputs and outputs, the company studies internal representations and the mechanisms that produce model behaviour.
Its latest research, published in September 2026, reports finding an internal signal associated with reward hacking and developing probes that can detect the behaviour at scale. The company says these monitors could allow developers to identify reward hacking during training and intervene before problematic behaviour becomes embedded in a model.

Goodfire Wants to Turn AI Interpretability Into an Engineering Discipline
Goodfire’s main product, Silico, is designed to make interpretability research more accessible and operational. It combines methods including sparse autoencoders, activation probing, causal analysis, data attribution, model comparison, neural geometry, and feature steering, while allowing researchers to run long-running experiments across models. The goal is to move from asking why a model produced an output to investigating which internal features and mechanisms contributed to that behaviour. For language models, Goodfire uses Silico to identify problematic behaviour, investigate reward hacking, build guardrails, and trace unexpected results back to training data or model behaviour.
Its research has also explored using internal model features as training signals. In one study, Goodfire reported reducing hallucinations in Google’s Gemma 3 12B by 58% using interpretability-derived rewards, without degradation on standard performance benchmarks. The company is extending the same approach beyond language models, including robotics and vision systems, where interpretability can help identify brittle representations and diagnose why physical AI systems fail.

From Detecting Reward Hacking to Designing AI From the Inside
Goodfire’s longer-term ambition is considerably broader than monitoring individual failures. The company describes its approach as an “intentional design” agenda, where interpretability becomes part of how AI systems are trained, debugged, and modified. Its research in life sciences illustrates what this could mean outside conventional AI safety. Goodfire says its tools have helped researchers identify biological signals inside foundation models, including a novel class of Alzheimer’s biomarkers, while its work with genomics models has focused on distinguishing genuine biological structure from dataset artefacts and other shortcuts.
In February 2026, Goodfire raised a $150 million Series B at a $1.25 billion valuation, bringing its total funding to more than $200 million according to the company. The round followed a $50 million Series A announced in 2025 and is supporting further research, product development, and partnerships across AI agents and life sciences.
The underlying bet is that increasingly capable AI will require more than better benchmarks and larger training runs. If developers can identify the internal concepts responsible for a model’s behaviour, they may gain a more precise way to understand failures, detect unwanted strategies, and deliberately shape what models learn. Reward hacking is one particularly visible example of the problem. The larger challenge is building enough understanding of neural networks that powerful AI systems can eventually be engineered with something closer to the precision of conventional software.

