Visual AI Digest: The Sora Rival Race Hits Prime Time

Multi · September 30, 2026 · 1 min read · 6 sources

News

Alibaba and ByteDance Launch Rival Text-to-Video Models

Alibaba and ByteDance have officially launched competing text-to-video models, directly challenging OpenAI's Sora. This intensifies the global race for accessible, high-quality AI video generation.

Research

Advancing Layout Control in Text-to-Image Generation

This paper introduces a new framework for generating images with precise spatial layout control from text prompts, a key challenge for practical T2I applications.

New Research on Subject Consistency in Long-Form Video Generation

Focuses on maintaining subject identity and consistency across multiple generated video frames, a major hurdle for creating usable AI video.

Reference-Based Motion Control for Text-to-Video Models

Explores methods for controlling motion dynamics in generated video using reference signals, moving beyond purely text-based prompts for more nuanced direction.

A Unified Transformer Architecture for Image and Video Synthesis

Presents a unified transformer architecture designed to handle both T2I and T2V tasks efficiently, potentially streamlining model development.

Improving Long-Prompt Fidelity in Text-to-Image Models

Tackles the problem of ensuring generated images faithfully align with very long and detailed textual descriptions, improving prompt fidelity.

Stay Ahead

Delivered each morning.