Skip to content
Signalcrest
← Back to feed
devtoai

Anthropic’s Reward Seeker Study Shows How Training Can Produce Misaligned AI Behavior

Study shows training can lead to misaligned AI behavior, warning developers.

Cooling

34

signal score

as of 1d ago

🔗 5

sources · claude

Too early

trajectory — needs a few more snapshots

Why this scored 34

every term, weighted
Velocity+0.0 / 40 max

Engagement gained per hour since the last capture, against the fastest item on its own source

Acceleration+4.0 / 25 max

Whether that velocity is itself speeding up, as a per-hour rate

Cross-source spread+20.0 / 25 max

How many independent communities are talking about the same entity

Recency+9.7 / 10 max

Decays to zero over 14 days

Saturation penalty0.0 / 30 max

Subtracted once something is big and old — we rank what's next, not what's peaked

Composite33.7

Weights are hand-tuned, not learned — we're calibrating them against realized trends as history accumulates. On a topic's first sighting there's no previous reading to compare against, so acceleration starts from a neutral prior rather than a measurement, and velocity falls back to engagement over its whole lifetime until a second reading exists. Full methodology

Signal history

7-day window (free)

Entities

anthropicclaude
Embed a live signal badge
Signalcrest signal badge
[![Signalcrest signal](https://www.signalcrest.app/api/badge/dev%3A4540863)](https://www.signalcrest.app/topic/dev%3A4540863)

Drop this in a README or blog post — it updates automatically as the score moves.