Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

ViDiC: Video Difference Captioning

PretrainZero: Reinforcement Active Pretraining






























papers

ViDiC: Video Difference Captioning

PretrainZero: Reinforcement Active Pretraining






























Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
SimScale: Learning to Drive via Real-World Simulation at Scale
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
Guided Self-Evolving LLMs with Minimal Human Supervision
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
How Far Are We from Genuinely Useful Deep Research Agents?
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
Mem-α: Learning Memory Construction via Reinforcement Learning
Search Self-play: Pushing the Frontier of Agent Capability without Supervision
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
Optimizing Mixture of Block Attention
FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
Chain-of-Thought Hijacking
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
3EED: Ground Everything Everywhere in 3D
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
CHIP: A multi-sensor dataset for 6D pose estimation of chairs in industrial settings
Geometrically-Constrained Agent for Spatial Reasoning
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
DiP: Taming Diffusion Models in Pixel Space
Architecture Decoupling Is Not All You Need For Unified Multimodal Model
Vision Bridge Transformer at Scale
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
SimScale: Learning to Drive via Real-World Simulation at Scale
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
Guided Self-Evolving LLMs with Minimal Human Supervision
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
How Far Are We from Genuinely Useful Deep Research Agents?
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
Mem-α: Learning Memory Construction via Reinforcement Learning
Search Self-play: Pushing the Frontier of Agent Capability without Supervision
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
Optimizing Mixture of Block Attention
FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks
Chain-of-Thought Hijacking
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
3EED: Ground Everything Everywhere in 3D
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
CHIP: A multi-sensor dataset for 6D pose estimation of chairs in industrial settings
Geometrically-Constrained Agent for Spatial Reasoning
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
DiP: Taming Diffusion Models in Pixel Space
Architecture Decoupling Is Not All You Need For Unified Multimodal Model
Vision Bridge Transformer at Scale
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement