Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

ACE-Step: A Step Towards Music Generation Foundation Model

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model






























papers

ACE-Step: A Step Towards Music Generation Foundation Model

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model






























Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?
A Flexible and Secure Deployment Framework for Distributed Applications
Multimodal Pretraining and Generation for Recommendation: A Tutorial
A Theoretical Limit to Physicalism: A Non-Technical Explanation of the Gemini Theorem
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
Propagation dynamics of the circular Airy Gaussian vortex beams in the fractional nonlinear Schrödinger equation
VASP on a GPU: application to exact-exchange calculations of the stability of elemental boron
Information quantity in a pixel of digital image
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
Recognition of Handwritten Roman Script Using Tesseract Open Source OCR Engine
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models
TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
Qwen2.5-Omni Technical Report
Dual-Scale Single Image Dehazing Via Neural Augmentation
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
DensityTool: A post-processing tool for space and spin-resolved density of states from VASP
A One-Dimensional Energy Balance Model Parameterization for the Formation of CO2 Ice on the Surfaces of Eccentric Extrasolar Planets
Towards The Ultimate Brain: Exploring Scientific Discovery with ChatGPT AI
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
MatterGen: a generative model for inorganic materials design
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
Phi-4 Technical Report
A Set of Tutorials for the LAMMPS Simulation Package
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?
A Flexible and Secure Deployment Framework for Distributed Applications
Multimodal Pretraining and Generation for Recommendation: A Tutorial
A Theoretical Limit to Physicalism: A Non-Technical Explanation of the Gemini Theorem
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
Propagation dynamics of the circular Airy Gaussian vortex beams in the fractional nonlinear Schrödinger equation
VASP on a GPU: application to exact-exchange calculations of the stability of elemental boron
Information quantity in a pixel of digital image
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
Recognition of Handwritten Roman Script Using Tesseract Open Source OCR Engine
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
Learning Temporal Evolution of Spatial Dependence with Generalized Spatiotemporal Gaussian Process Models
TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
Qwen2.5-Omni Technical Report
Dual-Scale Single Image Dehazing Via Neural Augmentation
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
DensityTool: A post-processing tool for space and spin-resolved density of states from VASP
A One-Dimensional Energy Balance Model Parameterization for the Formation of CO2 Ice on the Surfaces of Eccentric Extrasolar Planets
Towards The Ultimate Brain: Exploring Scientific Discovery with ChatGPT AI
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
MatterGen: a generative model for inorganic materials design
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
Phi-4 Technical Report
A Set of Tutorials for the LAMMPS Simulation Package
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot