Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

DarwinX: Evolving Agent Harnesses Through Natural Selection

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation






























Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
LLMROUTER: UNIFIED INFRASTRUCTURE FOR DEVELOPING, EVALUATING, AND DEPLOYING LLM ROUTERS
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
WizardLM: EMPOWERING LARGE PRE-TRAINED LANGUAGE MODELS TO FOLLOW COMPLEX INSTRUCTIONS
Indoor Segmentation and Support Inference from RGBD Images
Evaluating Language Models for Harmful Manipulation
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
MENDEL GÖDEL MACHINE: RECURSIVE SELF-IMPROVING CODING AGENTS VIA COMPARATIVE EVOLUTION
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
DeepTutor: Towards Agentic Personalized Tutoring
Flow-by-Flow: Content-Judgment Bypass for Governing AI Output in High-Loss Domains
Towards an Argumentative Foundation for Evaluative AI
Data-Driven Fire-Zone Segmentation for Improved Short-Term Wildfire Prediction
Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Skaling: Chinchilla’s Exponents Meet Kaplan’s Coupling
PROTECT-90: A Fault Dataset for Power System Protection
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family