AI

AI Pipeline Astra Converts Keynote Footage Into Editable 3D Blender Scenes

A demonstration using Apple executive Craig Federighi showcases how generative systems are moving from flat video outputs to parametric 3D reconstruction.

  • A newly demonstrated artificial intelligence workflow named Astra has managed to turn standard 2D presentation video into a fully editable 3D digital scene, signaling a significant shift in compute…
  • In a demonstration highlighted by @chatgptricks on Instagram, the system processed recorded presentation footage featuring Craig Federighi, senior vice president of software engineering at Apple.
  • Rather than outputting a flat, pixel-based generative clip, the pipeline creates structured assets: mesh geometry, camera vectors, and textured digital doubles.
AI Pipeline Astra Converts Keynote Footage Into Editable 3D Blender ScenesThe Scale Report

A newly demonstrated artificial intelligence workflow named Astra has managed to turn standard 2D presentation video into a fully editable 3D digital scene, signaling a significant shift in computer vision and automated content creation.

In a demonstration highlighted by @chatgptricks on Instagram, the system processed recorded presentation footage featuring Craig Federighi, senior vice president of software engineering at Apple. Using the reference clip as its sole guide, the model extracted the spatial layout, estimated camera paths, and reproduced both the background environment and the presenter directly inside the open-source 3D software Blender.

Rather than outputting a flat, pixel-based generative clip, the pipeline creates structured assets: mesh geometry, camera vectors, and textured digital doubles. While the reconstructed scene is not an identical frame-for-frame replica of the original broadcast, the system closely captures the structural composition, lighting, and movement of the original stage presentation.

Moving Past Pure Video Generation

Most high-profile generative video tools operate purely in pixel space, producing photorealistic renders that remain notoriously difficult to modify after generation. In contrast, converting video into native 3D assets provides artists and VFX supervisors with granular control over lighting rigs, camera focal lengths, and character meshes.

This shift matters because professional production pipelines in gaming, cinema, and virtual production rely heavily on editable scene graphs rather than flat rasterized files. If generative AI can reliably automate the tedious work of scene reconstruction and camera tracking, technical artists could dramatically cut pre-visualization and rotoscoping turnaround times.

Significant hurdles remain before automated scene extraction can match human-built production assets. Current iterations still exhibit geometry artifacts and simplified texturing, meaning studios will likely view these tools as rapid prototyping aids rather than complete replacements for traditional modeling and matchmoving workflows. Nonetheless, the ability to translate consumer-grade video into interactive 3D environments represents a clear step toward unifying generative models with traditional 3D software suites.

Reporting based on coverage from @chatgptricks on Instagram.

The daily brief

The biggest stories in AI, venture, sports business and culture - once a day.

One short email from The Scale Report. No spam, unsubscribe any time.

Read next