LPM1.0 Model Generates Real-Time Interactive Digital Human Video from Single Image

Researchers have officially unveiled the LPM1.0 model, a project designed to generate real-time videos of people speaking, listening, and singing from a single reference image. Its key breakthrough lies in multimodal processing—synchronizing text, audio, and image inputs to produce dynamic scenes with accurate lip sync, nuanced facial expressions, and natural emotional transitions. The model integrates directly with leading voice AI platforms like ChatGPT and Doubao, transforming conventional voice conversations into real-time interactive experiences with visual feedback.
Technically, LPM1.0 employs "multi-granularity identity conditioning" to extract detailed features from reference materials across multiple angles and expressions, eliminating the need for the model to independently generate complex elements like teeth, wrinkles, or side profiles. This greatly boosts cross-style processing—enabling instant drive for photorealistic faces, animations, or 3D game characters without retraining. The model also supports streaming transmission, ensuring system stability even when generating videos up to 45 minutes long.
Regarding interaction logic, LPM1.0 accurately detects three conversational states: while listening, it produces responsive expressions like nodding or shifting gaze; while speaking, it drives body and lip movements based on audio; while idle, it generates natural resting behaviors according to text prompts. Project manager Zeng Ailing noted that LPM1.0 is suitable not only for real-time conversation but also for offline audio-driven video generation, offering technical redundancy for podcasts and film production.
Despite its strong application potential, the development team stresses that LPM1.0 remains a research project with no current plans to release code or weights publicly. Researchers acknowledge a qualitative gap between generated videos and real footage, and the inherent deepfake risks cannot be overlooked. The study's significance lies in charting the future evolution of AI systems—shifting from single logical interaction to multidimensional interaction incorporating emotional responses, eye contact, and visual embodiment.
Related article
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Related Special Topic Recommendations
Comments (0)
0/500

Researchers have officially unveiled the LPM1.0 model, a project designed to generate real-time videos of people speaking, listening, and singing from a single reference image. Its key breakthrough lies in multimodal processing—synchronizing text, audio, and image inputs to produce dynamic scenes with accurate lip sync, nuanced facial expressions, and natural emotional transitions. The model integrates directly with leading voice AI platforms like ChatGPT and Doubao, transforming conventional voice conversations into real-time interactive experiences with visual feedback.
Technically, LPM1.0 employs "multi-granularity identity conditioning" to extract detailed features from reference materials across multiple angles and expressions, eliminating the need for the model to independently generate complex elements like teeth, wrinkles, or side profiles. This greatly boosts cross-style processing—enabling instant drive for photorealistic faces, animations, or 3D game characters without retraining. The model also supports streaming transmission, ensuring system stability even when generating videos up to 45 minutes long.
Regarding interaction logic, LPM1.0 accurately detects three conversational states: while listening, it produces responsive expressions like nodding or shifting gaze; while speaking, it drives body and lip movements based on audio; while idle, it generates natural resting behaviors according to text prompts. Project manager Zeng Ailing noted that LPM1.0 is suitable not only for real-time conversation but also for offline audio-driven video generation, offering technical redundancy for podcasts and film production.
Despite its strong application potential, the development team stresses that LPM1.0 remains a research project with no current plans to release code or weights publicly. Researchers acknowledge a qualitative gap between generated videos and real footage, and the inherent deepfake risks cannot be overlooked. The study's significance lies in charting the future evolution of AI systems—shifting from single logical interaction to multidimensional interaction incorporating emotional responses, eye contact, and visual embodiment.
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary





Home






