Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

Energy-Based Transformers are Scalable Learners and Thinkers






























papers

DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

Energy-Based Transformers are Scalable Learners and Thinkers






























IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
WebSailor: Navigating Super-human Reasoning for Web Agent
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Depth Anything at Any Condition
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
Kwai Keye-VL Technical Report
A Survey on Vision-Language-Action Models for Autonomous Driving
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks
Holistic Artificial Intelligence in Medicine; improved performance and explainability
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Listener-Rewarded Thinking in VLMs for Image Preferences
Calligrapher: Freestyle Text Image Customization
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
WebSailor: Navigating Super-human Reasoning for Web Agent
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Depth Anything at Any Condition
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
Kwai Keye-VL Technical Report
A Survey on Vision-Language-Action Models for Autonomous Driving
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks
Holistic Artificial Intelligence in Medicine; improved performance and explainability
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Listener-Rewarded Thinking in VLMs for Image Preferences
Calligrapher: Freestyle Text Image Customization
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models