Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data






























papers

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data






























Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
SpikingBrain Technical Report: Spiking Brain-inspired Large Models
SAGE: A Realistic Benchmark for Semantic Understanding
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
EmbeddingGemma: Powerful and Lightweight Text Representations
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
SIM-CoT: Supervised Implicit Chain-of-Thought
SWE-QA: Can Language Models Answer Repository-level Code Questions?
Video models are zero-shot learners and reasoners
An N-Plus-1 GPT Agency for Critical Solution of Mechanical Engineering Analysis Problems
Memory-QA: Answering Recall Questions Based on Multimodal Memories
MAPO: Mixed Advantage Policy Optimization
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
Reinforcement Learning on Pre-Training Data
Do You Need Proprioceptive States in Visuomotor Policies?
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
GenExam: A Multidisciplinary Text-to-Image Exam
Nav-R1: Reasoning and Navigation in Embodied Scenes
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
ARE: Scaling Up Agent Environments and Evaluations
DiffusionNFT: Online Diffusion Reinforcement with Forward Process
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
LIMI: Less is More for Agency
A Modular Fusion Neural Network Approach to Efficiently Predict Multi-Metal Binding Sites in Protein Sequences
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
SpikingBrain Technical Report: Spiking Brain-inspired Large Models
SAGE: A Realistic Benchmark for Semantic Understanding
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
EmbeddingGemma: Powerful and Lightweight Text Representations
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
SIM-CoT: Supervised Implicit Chain-of-Thought
SWE-QA: Can Language Models Answer Repository-level Code Questions?
Video models are zero-shot learners and reasoners
An N-Plus-1 GPT Agency for Critical Solution of Mechanical Engineering Analysis Problems
Memory-QA: Answering Recall Questions Based on Multimodal Memories
MAPO: Mixed Advantage Policy Optimization
Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
Reinforcement Learning on Pre-Training Data
Do You Need Proprioceptive States in Visuomotor Policies?
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
GenExam: A Multidisciplinary Text-to-Image Exam
Nav-R1: Reasoning and Navigation in Embodied Scenes
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
ARE: Scaling Up Agent Environments and Evaluations
DiffusionNFT: Online Diffusion Reinforcement with Forward Process
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
LIMI: Less is More for Agency
A Modular Fusion Neural Network Approach to Efficiently Predict Multi-Metal Binding Sites in Protein Sequences
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech