Skip to content
Signalcrest
Back to feed
devtoai

My Agent's Tests Were Green Because the Model Learned to Cheat

LLMs can game test suites, so developers must design robust evaluation methods.

Steady
Signal score
18
Trajectory
Cooling
7-day est. ~0

Why this scored 18

every term, weighted
Velocity+0.0 / 40

Engagement gained per hour since the last capture, against the fastest item on its own source

Acceleration+11.1 / 25

Whether that velocity is itself speeding up, as a per-hour rate

Cross-source spread+0.0 / 25

How many independent communities are talking about the same entity

Recency+8.2 / 10

Decays to zero over 14 days

Saturation penalty1.6 / 30

Subtracted once something is big and old — we rank what's next, not what's peaked

Composite17.7

Weights are hand-tuned, not learned — we're calibrating them against realized trends as history accumulates. On a topic's first sighting there's no previous reading to compare against, so acceleration starts from a neutral prior rather than a measurement, and velocity falls back to engagement over its whole lifetime until a second reading exists. Full methodology

Outlook

low confidence · estimate, not a guarantee

7-day

~0

range 027

14-day

~0

range 033

30-day

~0

range 044

Signal history

7-day window (free)
063

projected trajectory (estimate, not a guarantee)

Embed a live signal badge
Signalcrest signal badge
[![Signalcrest signal](https://www.signalcrest.app/api/badge/dev%3A4654117)](https://www.signalcrest.app/topic/dev%3A4654117)

Drop this in a README or blog post — it updates automatically as the score moves.