Archive
Everything we've scored that's past the 48-hour window — open to everyone, no account needed. The live feed only shows what's moving right now; this is the record behind it.
- 19BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing
- 19Fusion Training for Mathematical Generalization in Large Language Models
- 19Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
- 19Beyond Local Surprise: Grounded Dialogue as Selective Belief Revision under Referential Uncertainty
- 19Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions
- 19Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
- 19Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
- 19Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement
- 19Aspire: Can Models Self-Evolve from Vague Goals?
- 19Consilience for Verifier-Free Test-Time Scaling
- 19SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- 19What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
- 19MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents
- 19When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay
- 19Fairness in Link Prediction Beyond Demographic Parity: A Reproducibility Study
- 19rustgrep - structural grep for Rust source
- 19Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
- 19Parameterized Complexity of $L_p$-Lipschitz Constants for Input Convex Neural Networks and $L_p$-Norm Maximization over Zonotopes
- 19zLend: A Dual-Scope Cash-Flow Reconstruction Framework for On-Chain Credit Underwriting
- 19DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
- 19DSLE: A Learning Environment for Dark Souls Boss Encounters
- 19The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
- 19Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity
- 19Designing Proactive Thought Partners for Writing
- 19Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
- 19StudentSim: Training LLM-based Student Simulators
- 19Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
- 19A Generalization of Amari's Bayesian Duality
- 19A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
- 19SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- 19EIP-8379: Top-up Sync
- 19The Value of Human Expertise
- 19How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention
- 19$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
- 19The canonical facets of multi-separator polytopes
- 19Mechanism Design for Alignment and Control
- 19Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks
- 19Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
- 19The Rise of Verbal Reinforcement Learning
- 19ExecCritic: Learn to Test, Test to Improve for Coding Agents
- 19Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
- 19Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation
- 19CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
- 19Non-Crossing Deep Quantile Regression for Distributional Survival Prediction
- 19TimescaleDB 2.27 Added Bloom Filters to UPDATE and DELETE. Your EXPLAIN Won't Tell You If They Work Unless You Know These Counters.
- 19Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
- 19Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
- 19Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
- 19A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes
- 19"Train classical, deploy quantum" requires rethinking generalization
- 19Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
- 19NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
- 19Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
- 19GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
- 19Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
- 19Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
- 19Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
- 19What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
- 19From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
- 19Multimodal Model Diffing for Feature Discovery and Control
- 19Redistribution-based Cost Inference Improves Sparse Safe Offline RL
- 19AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- 19Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
- 19When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning
- 19Copying explains the collective behavior of AI agents in the wild
- 19Silver Rate Is (Almost) Optimal for Gradient Descent Acceleration
- 19Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
- 19Data-Efficient and Interpretable Classification of Circulating Tumor Cell Phenotypes in Microfluidic Devices via Deep Learning
- 19ReCite: Agentic Reasoning for Faithful Citation
- 19Learning Length-Extrapolatable Recurrent Models
- 19TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
- 19DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
- 19PaperGym: Rubric-Centered Evolution for Research-Plan Generation
- 19EIP-8372: Normalized state gas limit
- 19An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators
- 19On the Complexity of the Compatibility Problem for Succinctly Encoded Conditional Distributions
- 19AutoSR: Automatic Symbolic Regression by Searching Research States
- 19AVA-Encoder: Towards Agent-Native Video Representation Learning
- 19Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs
- 19Prefix Sliding for efficient test-time scaling
- 19Spectral Gaps of Hit-and-Run and Coordinate Hit-and-Run
- 19Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
- 19Q-based Variational Inverse Reinforcement Learning
- 19Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
- 19EIP-8371: RowDAS - Distributed Blobspace Reconstruction
- 19Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving
- 19Show HN: Open-source Stripe Connect alternative
- 19DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening
- 19A rule-based Korean dialect converter with no LLM: eight guards that keep standard Korean intact
- 19Implementing neural network mixed-effects models in Template Model Builder (TMB)
- 19SwarmWorld: Stigmergic technological evolution in societies of language-model agents
- 19OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques
- 19Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation
- 19You still write `finally { resource.close() }`. The `using` keyword does it automatically.
- 19ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
- 19Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification
- 19TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
- 19Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
- 19Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
- 19PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
- 19Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role
- 19MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
- 19A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
- 19Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations
- 19All Core Devs - Testing (ACDT) #96, September 14, 2026
- 19VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
- 19FieldOS, Part 3: The App People Open, and What 10 Hours Taught Me About Pricing
- 19SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies
- 19Constant Individual Regret in General Games
- 19Context-Aware Interleaved Batching for WhisperX
- 19Doesn't JavaScript have console input?
- 19Four Debian 13 Boxes, One Brief: 1,923 Packages on Metal, 328 in the Cloud
- 19I shipped 200 tools as single HTML files. Here's what that constraint actually buys you.
- 19My AI Agent Tried to Delete Every Customer Record. Here's What Stopped It.
- 19AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
- 19Discord-Nitro-Generator — 🚀 Advanced Discord Nitro Generator and Checker (2026). Fast, multi-threaded, and proxy-supported free Discord Nit
- 19Your .mcp.json probably has a live API key in it
- 19Gleam has mirrored its source code on tangled (an AT-protocol based forge)
- 19Are we the abstraction? AI and the future of software engineering
- 19117 Ghost Errors: Anatomy of a Flaky AI Agent
- 19RPC Standards # 34, September 7th, 2026
- 19QGPINNs: A Physics-Informed Neural Network Framework for Nonlocal Differential Equations on Quantum Graphs
- 19Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning
- 19Learning a Size-Weight Frontier for Synthetic-Augmented Inference
- 19On two proofs of $d^2$ mixing of weighted Dikin walks
- 19Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy
- 19A Formal Limitation on Learning Human Language From Textual Corpora
- 19I analyzed 14,905 love letters without reading them
- 19Blog: Survey of Optimizers
- 19“We have information that Moonshot distilled Fable for the development of K3”
- 19ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
- 19Information on trajectories: martingales and random times
- 19G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
- 19Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
- 19Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
- 19ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
- 19AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
- 19OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
- 19Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
- 19Logos: An Agent Harness on a Cross-Process Bus
- 19One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
- 19HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
- 19Robust PAC Learning of Concurrent Stochastic Games
- 19$TCP_α$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval
- 19Defensive Boosting for Online Probabilistic Forecasting
- 19Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure
- 19A comparison between ceiling-mounted FMCW, IR-UWB and Wi-Fi radar for in-bedroom human activity monitoring and sleep interruption detection
- 19Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
- 19Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining
- 19An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction
- 19CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
- 19WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- 19Inducing Task Models from Computer-Use Traces
- 19AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- 19Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
- 19QuoteBench: How Matched Scores Can Hide Command-Path Failures
- 19TTPO: Test-Time Policy Optimization
- 19SWE-Prime: Fewer Trajectories, Better Performance
- 19A Computationally Feasible Framework for Causal Probabilistic Explanation
- 19Extending Slugs Across Templates and Entities for Deterministic API Workflows
- 19LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
- 19@llm-ports/core — Provider-agnostic LLM port interface, registry, cost gating, content blocks, validation strategies. The foundation of llm-
- 19Last Translation Benchmark
- 19Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
- 19Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- 19From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
- 19Video Generative Models as Geometry Learner
- 19Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records
- 19Using df.styler with df.to_latex() and longtable
- 19A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
- 19MidTool: Mid-training Data Synthesis for Agentic Tool Use
- 19Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs
- 19RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
- 19SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
- 19SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
- 19From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
- 19DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging
- 19Parameterised graph theory for tensor networks: entanglement rerouting, structural simplification, and agnostic tomography
- 19SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
- 19Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
- 19Stochastic Estimation of Transduced Language Models
- 19Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
- 19Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
- 19Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils
- 19Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
- 19An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models
- 19Boosting LLM Exploration via Weak-Model Guidance in RLVR
- 19Marionette: Predicting World States, Rendering Geometry, Painting Appearance
- 19DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
- 19Handover of In-Context Learning State Across Session Boundaries
- 19Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
- 19Vero: Can AI Agents Build Formally Verified Software Repositories?
- 19A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle
- 19Exponential quantum advantage for learning signals with a single qubit
- 19Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
- 19The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity
- 19Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers
- 19InstructMesh: Selective Refinement of Generative 3D Models for Fabrication
- 19Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology
- 19DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data