ByteDance and HKUST Unveil MMProLong to Boost Long Document LMM Training

On May 24, the ByteDance Seed team partnered with the Hong Kong University of Science and Technology to unveil a breakthrough in long-document training for multimodal large language models (LMMs). Leveraging Alibaba’s open-source Qwen2.5-VL, the team developed a new model named MMProLong, achieving significant gains in processing efficiency. This research challenges conventional long-text training methods for multimodal models, highlighting how strategic data organization is critical to enhancing long-context capabilities.
Addressing key challenges in LMM training, the study reveals that question-and-answer (QA) training tailored to specific tasks is far more effective than traditional optical character recognition (OCR) transcription. While relying solely on text transcription fails to improve content retrieval in long contexts—and can even degrade performance—training with long-context QA pairs generated by an independent model like ByteDance Seed2.0 enables the model to accurately locate target paragraphs amidst lengthy, distracting information.
With an optimized strategy and a modest training budget of just 128,000 tokens, MMProLong maintains robust long-text stability, performing effectively even when input lengths reach 256,000 or 512,000 tokens. It surpasses larger open-source models such as InternVL3-38B and Gemma3-27B on the MMLongBench and MM-NIAH (Needle-in-a-Haystack) benchmarks. Furthermore, MMProLong’s multimodal capabilities extend to long-video understanding tasks without specific training, a strategy also validated on the Qwen3-VL-8B model.
This study offers an alternative development path for the large model industry, contrasting with approaches like DeepSeek’s, which focus on upgrading architectures through highly compressed and re-ordered visual data. It demonstrates that long-context capabilities can be significantly enhanced by optimizing training data structure rather than altering the underlying architecture, paving the way for more cost-effective and efficient technical solutions for future multimodal systems and multi-step agents.
Related article
Microsoft revives in-house AI strategy with new coding model as Claude costs soar
AI-assisted coding has become a staple in modern software development, yet even tech giants like Microsoft are feeling the pinch of expensive third-party large language model subscriptions. To cut external dependencies and operational overhead, Micro
Another customer of struggling startup Delve hit by major security breach
The compliance startup Delve continues to face a series of escalating controversies.TechCrunch has verified that Delve conducted the security certifications for Context AI, an AI agent training firm that recently reported a security incident resultin
Google Unveils Gemini 3.8 Flash TTS, Enabling Natural Language Voice Customization and 30-Second Replication
Speech synthesis is reaching new heights of precision. On September 23, Google unveiled two advanced text-to-speech models—Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS—designed to meet the diverse demands of developers and creators for personal
Related Special Topic Recommendations
Comments (0)
0/500

On May 24, the ByteDance Seed team partnered with the Hong Kong University of Science and Technology to unveil a breakthrough in long-document training for multimodal large language models (LMMs). Leveraging Alibaba’s open-source Qwen2.5-VL, the team developed a new model named MMProLong, achieving significant gains in processing efficiency. This research challenges conventional long-text training methods for multimodal models, highlighting how strategic data organization is critical to enhancing long-context capabilities.
Addressing key challenges in LMM training, the study reveals that question-and-answer (QA) training tailored to specific tasks is far more effective than traditional optical character recognition (OCR) transcription. While relying solely on text transcription fails to improve content retrieval in long contexts—and can even degrade performance—training with long-context QA pairs generated by an independent model like
With an optimized strategy and a modest training budget of just 128,000 tokens, MMProLong maintains robust long-text stability, performing effectively even when input lengths reach 256,000 or 512,000 tokens. It surpasses larger open-source models such as InternVL3-38B and
This study offers an alternative development path for the large model industry, contrasting with approaches like DeepSeek’s, which focus on upgrading architectures through highly compressed and re-ordered visual data. It demonstrates that long-context capabilities can be significantly enhanced by optimizing training data structure rather than altering the underlying architecture, paving the way for more cost-effective and efficient technical solutions for future multimodal systems and multi-step agents.
Microsoft revives in-house AI strategy with new coding model as Claude costs soar
AI-assisted coding has become a staple in modern software development, yet even tech giants like Microsoft are feeling the pinch of expensive third-party large language model subscriptions. To cut external dependencies and operational overhead, Micro
Another customer of struggling startup Delve hit by major security breach
The compliance startup Delve continues to face a series of escalating controversies.TechCrunch has verified that Delve conducted the security certifications for Context AI, an AI agent training firm that recently reported a security incident resultin
Google Unveils Gemini 3.8 Flash TTS, Enabling Natural Language Voice Customization and 30-Second Replication
Speech synthesis is reaching new heights of precision. On September 23, Google unveiled two advanced text-to-speech models—Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS—designed to meet the diverse demands of developers and creators for personal





Home






