rl
for ai researchers
Why this scored 0
every term, weightedIts fastest-moving item, measured against the pace of its own source
Whether that velocity is itself speeding up, as a per-hour rate
How many independent communities its own items come from
Decays to zero over 14 days, counted from when we first saw it
Subtracted once something is big and old — sized by its biggest item, aged from when we first saw it
Weights are hand-tuned, not learned — we're calibrating them against realized trends as history accumulates. On an entity's first sighting there's no previous reading to compare against, so acceleration starts from a neutral prior rather than a measurement, and velocity falls back to engagement over its whole lifetime until a second reading exists. Full methodology
Signal history
7-day window (free)The evidence
The live items this entity's score aggregates — every community independently talking about it right now. This is the corroboration, shown, not claimed.
- 112
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
Pretraining Q-functions · for ai researchers
arxivSteady 2 · rlai43d ago - 26
A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
Fine-tuned model beats frontier models · for ai researchers
hackernewsSteadyai45d ago