Visual AI Digest: Sora Challengers, Physical Acoustics, and 3D Scene Speed

Multi · September 16, 2026 · 1 min read · 6 sources

News

Alibaba and ByteDance Drop New Models to Challenge Sora

Alibaba and ByteDance are taking the fight directly to OpenAI with high-capability video models. This signals a major shift where Asia-Pacific tech giants are driving the narrative in high-fidelity text-to-video generation.

Research

Bounding Box Control for High-Fidelity Text-to-Video

Grounding T2V generation against specific bounding boxes and trajectories is notoriously difficult. This architecture proposes a cleaner separation of object layout and temporal motion, which is critical for practical editing workflows.

Solving Slot-Conflicting in Text-to-Image Generation

Most T2I models struggle to maintain distinctness in scenes with more than three objects. This paper introduces a new slot-conflict resolution mechanism that prioritizes layout grounding over text attention.

Fast Unsupervised 3D Scene Reconstruction from Text

View synthesis is getting faster without the usual texture blurring hazards. By using a

Physical Acoustics in Text-to-Audio Generation

Text-to-audio generation often ignores how objects interact physically. This study grounds sound generation in physical properties—like gravity and material density—making audio-visual pairs much more coherent.

Zero-Shot Training-Free Subject-Driven Video Generation

Maintaining temporal consistency in long video edits often requires expensive training. This approach introduces a frictionless pipeline for subject-driven video generation that requires zero additional training.

Stay Ahead

Delivered each morning.