Archive
Everything we've scored that's past the 48-hour window — open to everyone, no account needed. The live feed only shows what's moving right now; this is the record behind it.
- 19HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression
- 19Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
- 19Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension
- 19Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
- 19Automatic Model Card Generation Using an LLM
- 19GRADSOLVE: fast exact gradients for ODE ensembles on GPUs
- 19Graph Machine: Towards Better Pretraining via Edges
- 19EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
- 19Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule
- 19Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch
- 19RA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classification
- 19Traceable Trust for action-ready artificial intelligence in bioscience
- 19Sierpiński--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein Approximation
- 19Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core
- 19VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
- 19Mismatch Matters: On-Policy Distillation Beyond Token Agreement
- 19Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
- 19Discriminative World Models for Web Agents
- 19A Common Measure of Communication for Speech Brain-Computer Interfaces
- 19Discretizing Continuous Time Series for Imputation with Masked Diffusion Training
- 19An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
- 19PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
- 19Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
- 19Your Agent Isn't Losing Memory. It's Rotting.
- 19Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
- 19MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification
- 19MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
- 19Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering
- 19Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
- 19Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
- 19From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
- 19Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting
- 19Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles
- 19A Geometric Theory of Robust Fairness Audits
- 19Neurosymbolic Embodied Agents
- 19TabNSM: Neural Sparse Mixer for Tabular Regression
- 19The Measurement Revolution? Credible Measurement and Inference in the Age of AI
- 19Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
- 19Chain-of-Experience for Continual LLM Improvement
- 19Can LLMs Design Video Coding Tools? A Case Study on Planar Mode
- 19Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis
- 19Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields
- 19A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex
- 19EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
- 19Correcting a learned physical invariant improves world-model rollouts
- 19Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval
- 19UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation
- 19ClawGym II: Exploring Black-Box RL on Agent Harness
- 19I built a PDF table checker that verifies arithmetic in any PDF without a template
- 19One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
- 19BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval
- 19Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
- 19Kill The Cookie Banner
- 19Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement
- 19Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry
- 19Real-Time Climate Risk Assessment for Supply Chain Resilience: A Data-Driven Nowcasting Framework for Colombian Agriculture
- 19Comment-level Topic Drift Analysis in the Reddit Corpus
- 19SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval
- 19Variable Selection for Feature-Based Newsvendor
- 19When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
- 19A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
- 19CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
- 19Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
- 19Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
- 19Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors
- 19Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting
- 19Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
- 19RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
- 19Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
- 19How to Verify Consistency of Probabilistic Claims
- 19ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs
- 19Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization
- 19Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
- 19SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue
- 19Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
- 19When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
- 19NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games
- 19Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
- 19Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation
- 19A Mathematical Theory of Reusable Neural Bases for Network Compression
- 19Unsupervised Learning of Cell Instances with Generative Routing Pyramids
- 19Pretraining Reusable Inference Across Views with Synthetic Task Priors
- 19Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition
- 19Agentic Auto-Research is Fuzz Testing
- 19Quipu: A Governed Bitemporal Knowledge Graph Store
- 19Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
- 19Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores
- 19Interpretable AI with Local Distillation
- 19BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net
- 19A Model with No Head and Many Thoughts
- 19Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
- 19Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing
- 19Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
- 19Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks
- 19A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement
- 19CardioFusion-AI: Robust ECG--PPG Fusion for Multimodal Physiological Monitoring Under Signal Degradation
- 19The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams
- 19Enhancing EBSD throughput of battery electrode materials using super-resolution generative adversarial networks
- 19Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring
- 19H3-World: Turning Language Understanding into World Control
- 19Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling
- 19How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles
- 19AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
- 19Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling
- 19VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
- 19Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- 19Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions
- 19Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
- 19Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
- 19Towards Expert-level Medical AI for Real-time Video Consultations
- 19Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
- 19A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
- 19A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery
- 19Inertial Manifold Neural Operator for Dissipative Time-Dependent Partial Differential Equations
- 19GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
- 19Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination
- 19Imitation Learning for Connection-Tableau Construction
- 19Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication
- 19StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
- 19Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization
- 19Interpretable AI predicts a 2026 summer dry anomaly in central China
- 19Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration
- 19Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
- 19ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos
- 19Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
- 19CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
- 19VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
- 19Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
- 19Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training
- 19Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
- 19How My AI Agents' Mistakes Become Permanent Rules (And Why I Want Them to Fail)
- 19Stealing Reasoning Traces from Proprietary LLM APIs
- 19Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms
- 19Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction
- 19Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
- 19ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings
- 19ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
- 19Prime Agent: A Self-Improving RLM Harness
- 19The First Token Is a Clue: Verbalizing Multi-Token Concepts from the J-lens
- 19A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains
- 19Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
- 19Provably adaptive sampling with uniform and remasking discrete diffusion models
- 19HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance
- 19From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification
- 19Time-Aware Validation of Machine Learning Fuel Consumption Models: Evidence from 1\,Hz Operational Data, CCGS \textit{Sir Wilfrid Laurier}
- 19ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback
- 19Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
- 19Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
- 19ArchAgent v2: A Case Study with the Data Prefetching Championship
- 19Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention
- 19Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
- 19Model Hypnosis: Strong control of AI via additive subliminal effects
- 19Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
- 19DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation
- 19The concentration game: Bayesian updating, regret, and information
- 19Minimax bounds for watermarked and masked recursive discrete distribution estimation
- 19Finetuning Strategies for Querying Sounds by Vocal Imitation
- 19TokEval: A Tokenizer Evaluation Suite
- 19Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- 19EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings
- 19SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
- 19ReWorld: An Interactive World Model with Long-Horizon Memory
- 19ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
- 19ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
- 19Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
- 19How to Train a Critic Stably and Efficiently
- 19HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
- 19Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
- 19I Found a Hidden Layer Inside AI Images
- 19Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
- 19On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
- 19GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting
- 19Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
- 19Financial Numerical Prediction and Allocation as Token Generation
- 19ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
- 19It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention
- 19One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual Learning
- 19Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations
- 19VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
- 19Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation
- 19LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
- 19FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
- 19Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals
- 19SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
- 19From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
- 19S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
- 19Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
- 19Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
- 19BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes
- 19From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
- 19Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics
- 19SPADE: Self-Play in Adaptive Synthetic Executable Environments
- 19Form-Associated Custom Elements: Web Components That Belong in a Form
- 19The Surprising Effectiveness of Approximate Value Iteration in Self-Play
- 19LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
- 19Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football
- 19BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
- 19Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
- 19Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics
- 19BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing