Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

Agentopia: Long-Term Life Simulation and Learning in Agent Societies






























MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
GAN-BERT: Generative Adversarial Learning for Robust Text Classification with a Bunch of Labeled Examples
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
AVA-Encoder: Towards Agent-Native Video Representation Learning
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
ASI-BENCH: AT THE DAWN OF ARTIFICIAL SUPERINTELLIGENCE
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements
Demystifying Agent Skills: Why They Work—Until They Don’t
Open Mathematical Problems as an AI Reasoning Benchmark
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Autonomous de novo protein binder design with Claude
Agentic Transaction: Towards ACID-Compliant Agent Systems
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
ClawGym II: Exploring Black-Box RL on Agent Harness
MOSS-VL Technical Report
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
VIBEWORLDING: CAN MULTIMODAL AGENTS CONSTRUCT 3D OPEN WORLDS END-TO-END?
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation
Adversarial Learning of Classifier-Free Guidance Schedules
Agent-Orchestration in Autonomous Chip Design
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
Demonstration of Space Robot Teleoperation over a Lossy and Delayed Network using ATMOS
CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution
Training AI Scientists to Replicate Research
Small-Scale Experiments: Are We There Yet?
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Intern-S2-Preview: Scientific Agentic Foundation Model