Video Generation with AI: Working with Diffusion Transformers and Multimodal Learning

Video Generation with AI: Working with Diffusion Transformers and Multimodal Learning book cover

Video Generation with AI: Working with Diffusion Transformers and Multimodal Learning

Author(s): Joseph Enochs (Author)

  • Publisher: O’Reilly Media
  • Publication Date: August 25, 2026
  • Edition: 1st
  • Language: English
  • Print length: 294 pages
  • ASIN: B0GKF2TZ39
  • ISBN-13: 9798341653344

Book Description

Video generation is rapidly becoming a key area in generative AI, combining spatial, temporal, and multimodal reasoning to produce moving images that are both coherent and creative. For many practitioners, however, understanding how these models function and implementing them remains a significant challenge. Video Generation with AI offers a straightforward guide for exploring this new terrain.

Author Joseph Enochs leverages his experience leading enterprise AI projects to clarify how diffusion transformers, multimodal large language models, and spatiotemporal architectures combine to create high-quality video. Blending technical detail with practical examples, he demonstrates how to transition from isolated experiments to production-ready systems that transform creative fields, media, and human-machine collaboration.

  • Understand the fundamental architectures of modern video generative models
  • Train and fine-tune models using diffusion transformers and multimodal encoders
  • Maintain temporal coherence and consistency across complex scenes and sequences
  • Apply generative video tools to creative, industrial, and scientific workflows
  • Evaluate, troubleshoot, and improve generated video quality at scale

Editorial Reviews

Editorial Reviews

About the Author

Joseph Enochs is a technology executive, AI strategist, and award-winning speaker with over 15 years of experience in driving transformative Enterprise Data and AI solutions. He currently serves as Managing Director of AI/ML and Emerging Technologies at EVT, where he pioneers the development of accessible and ethical Generative AI systems. Joseph has delivered keynote presentations on AI’s role in enterprise transformation, most recently for The Profound Podcast and Disney’s Internal JETA Training Academy.

View on Amazon

未经允许不得转载:Wow! eBook » Video Generation with AI: Working with Diffusion Transformers and Multimodal Learning