2 min read
Add as a preferred source on Google

Goodfire Says Internal AI Monitors Cut Rogue-Agent Detection Costs

The company says its probes detected 94% of malicious hacking sessions while adding less than 2% to response-start time.

Sensor probes arranged beside a compact testing computer / TokenPost.ai
Sensor probes arranged beside a compact testing computer / TokenPost.ai

Goodfire says its internal-activation monitors detected 94% of malicious hacking sessions in tests while costing substantially less than conventional AI systems that reread an agent’s outputs.

The monitors inspect signals inside an AI model as it works, allowing a separate system to be called only when a potential risk appears. Baseten customers can configure the technology to detect offensive hacking, chemical and biological weapons misuse, reward hacking and other risks.

Customers may configure the system to record a flagged incident, escalate it to a person or block the underlying request.

Goodfire said its tests involved about 1,500 sessions with the open Kimi K3 model. Monitoring those sessions cost roughly $51, compared with about $233 for a lower-cost AI monitor that reviewed each step and approximately $10,000 for a top-tier model.

The company said its probes detected 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second look. Running four probes increased the model’s response-start time by less than 2%.

“Internal activation monitors are really cheap because they reuse the computations in the forward pass. So the model’s already computing this token,” Goodfire CEO Eric Ho said. He said the system takes intermediate neural activations that the model has already calculated and runs a classifier over them.

The small detectors, known as probes, scan internal model signals at each stage of an agent’s work. A separate AI model conducts a closer review only after a probe identifies a possible threat.

Goodfire CTO and co-founder Dan Balsam said the approach can identify harmful behavior before it is completed, including while a model is being evaluated or trained.

The company is targeting open models that developers can download and modify, including by removing built-in safeguards. Goodfire’s research found that Kimi K3 and GLM 5.2 reward-hacked in 50% to 96% of agent tests.

Balsam said the monitors are part of a broader effort to trace model behavior back to how it developed during training.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…