Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

The Diffusion Duality

Effective Red-Teaming of Policy-Adherent Agents






























papers

The Diffusion Duality

Effective Red-Teaming of Policy-Adherent Agents






























Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
Unified differentiable learning of electric response
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
Text-Aware Image Restoration with Diffusion Models
Magistral
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
Sapiens: Foundation for Human Vision Models
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
SAM 2: Segment Anything in Images and Videos
The Llama 3 Herd of Models
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
What matters when building vision-language models?
DDOS: The Drone Depth and Obstacle Segmentation Dataset
Deep learning-based framework for the on-demand inverse design of metamaterials with arbitrary target band gap
PRefLexOR: preference-based recursive language modeling for exploratory optimization of reasoning and agentic thinking
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
PlayerOne: Egocentric World Simulator
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis
Efficient Machine Learning Force Field for Large-Scale Molecular Simulations of Organic Systems
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education
ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation
Sequence Model Design for Code Completion in the Modern IDE
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
Unified differentiable learning of electric response
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
Text-Aware Image Restoration with Diffusion Models
Magistral
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
Sapiens: Foundation for Human Vision Models
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
SAM 2: Segment Anything in Images and Videos
The Llama 3 Herd of Models
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
What matters when building vision-language models?
DDOS: The Drone Depth and Obstacle Segmentation Dataset
Deep learning-based framework for the on-demand inverse design of metamaterials with arbitrary target band gap
PRefLexOR: preference-based recursive language modeling for exploratory optimization of reasoning and agentic thinking
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
PlayerOne: Egocentric World Simulator
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis
Efficient Machine Learning Force Field for Large-Scale Molecular Simulations of Organic Systems
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education
ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation
Sequence Model Design for Code Completion in the Modern IDE