Daily Paper Cast
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected]
Creator:
Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/
Gengyu Wang, LLM ML, http://wanggengyu.com
Listen on:
Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL
Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236
Cover Image by Kawen Kuang https://kawen.art
Informazioni
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected]
Creator:
Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/
Gengyu Wang, LLM ML, http://wanggengyu.com
Listen on:
Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL
Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236
Cover Image by Kawen Kuang https://kawen.art
1453
Episodi
English
Lingua
0
Iscritti
0
Ascolti
Episodi25
TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
Dec 8, 2025·—
—
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
Dec 8, 2025·—
—
From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
Dec 8, 2025·—
—
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
Dec 8, 2025·—
—
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Dec 5, 2025·—
—
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
Dec 5, 2025·—
—
Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
Dec 5, 2025·—
—
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
Dec 5, 2025·—
—
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
Dec 5, 2025·—
—
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
Dec 5, 2025·—
—
PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing
Dec 5, 2025·—
—
Qwen3-VL Technical Report
Dec 4, 2025·—
—
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
Dec 4, 2025·—
—
PretrainZero: Reinforcement Active Pretraining
Dec 4, 2025·—
—
ViDiC: Video Difference Captioning
Dec 4, 2025·—
—
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Dec 3, 2025·—
—
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
Dec 3, 2025·—
—
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
Dec 3, 2025·—
—
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
Dec 3, 2025·—
—
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
Dec 3, 2025·—
—
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
Dec 3, 2025·—
—
Guided Self-Evolving LLMs with Minimal Human Supervision
Dec 3, 2025·—
—
SimScale: Learning to Drive via Real-World Simulation at Scale
Dec 3, 2025·—
—
InnoGym: Benchmarking the Innovation Potential of AI Agents
Dec 3, 2025·—
—
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Dec 2, 2025·—
—


