Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory






























papers

GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory






























Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving
Story2Board: A Training-Free Approach for Expressive Storyboard Generation
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
Llama-Nemotron: Efficient Reasoning Models
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
Virtual staining of label-free tissue in imaging mass spectrometry
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
CharacterShot: Controllable and Consistent 4D Character Animation
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
Matrix-3D: Omnidirectional Explorable 3D World Generation
WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent
Marco-Voice Technical Report
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
Design of highly functional genome editors by modelling CRISPR–Cas sequences
UserBench: An Interactive Gym Environment for User-Centric Agents
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
WideSearch: Benchmarking Agentic Broad Info-Seeking
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
AWorld: Dynamic Multi-Agent System with Stable Maneuvering for Robust GAIA Problem Solving
Story2Board: A Training-Free Approach for Expressive Storyboard Generation
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
Llama-Nemotron: Efficient Reasoning Models
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
Virtual staining of label-free tissue in imaging mass spectrometry
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
CharacterShot: Controllable and Consistent 4D Character Animation
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
Matrix-3D: Omnidirectional Explorable 3D World Generation
WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent
Marco-Voice Technical Report
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C
Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
Design of highly functional genome editors by modelling CRISPR–Cas sequences
UserBench: An Interactive Gym Environment for User-Centric Agents
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
WideSearch: Benchmarking Agentic Broad Info-Seeking
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
Adapting Vision-Language Models Without Labels: A Comprehensive Survey