Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving






























SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning Paradigms for LLMs
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
ENVACE: INTERNALIZING ENVIRONMENT DYNAMICS VIA WORLD REHEARSAL FOR AGENTIC REINFORCEMENT LEARNING
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
WorldClaw: Agentic 3D Open-World Generation at Scale
INTERPRETABLE MEG DECODING OF PERCEIVED SPEECH: CORTICAL SOURCES AND THE STIMULUS FEATURES THAT DRIVE RETRIEVAL
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
AGENTOPSD: RECURSIVE SELF-DISTILLATION FOR AGENTIC REINFORCEMENT LEARNING
ACM: Agentic Context Management for Long Horizon Tasks
AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean Shift
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
REUSING ROLLOUTS UNDER POLICY LAG: PREFIX-NORMALIZED POLICY OPTIMIZATION FOR LLM REIN-FORCEMENT LEARNING
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
FinanceHarness: Autonomous Financial Deep Research Framework
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Knowledge–Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
TRAINING nGPT