aj-concepts.net

Gen AI Development

I've used various generative techniques and agentic workflows on commercial projects, but all of them are still under NDA. These are my personal experiments in using Gen AI tools in something approaching actual production

I've used various generative techniques and agentic workflows on commercial projects, but all of them are still under NDA. These are my personal experiments in using Gen AI tools in something approaching actual production

For this test I wanted a fully controllable video-to-video workflow driven by a 3D scene, rather than the 'slot machine' of text-to-image-to-video. In an ideal world, this would be a final 'skinning' pass on top of a rough 3D render, adding the last 15% polish that commonly eats up the majority of the time

For this test I wanted a fully controllable video-to-video workflow driven by a 3D scene, rather than the 'slot machine' of text-to-image-to-video. In an ideal world, this would be a final 'skinning' pass on top of a rough 3D render, adding the last 15% polish that commonly eats up the majority of the time

This is a diagram of the full workflow on a custom 'Obsidian' canvas. The Obsidian canvas isn't just a presentation tool here - it's the interface I used to create, interact and document the process of creating a full environment replacement for an exisiting plate

This is a diagram of the full workflow on a custom 'Obsidian' canvas. The Obsidian canvas isn't just a presentation tool here - it's the interface I used to create, interact and document the process of creating a full environment replacement for an exisiting plate

I integrated a number of different cloud platforms for this work, including Magnific, Beeble and Slapshot. In the end, most generative passes were done directly by a Cursor agent through the relevant API
I integrated a number of different cloud platforms for this work, including Magnific, Beeble and Slapshot. In the end, most generative passes were done directly by a Cursor agent through the relevant API
I integrated a number of different cloud platforms for this work, including Magnific, Beeble and Slapshot. In the end, most generative passes were done directly by a Cursor agent through the relevant API
I integrated a number of different cloud platforms for this work, including Magnific, Beeble and Slapshot. In the end, most generative passes were done directly by a Cursor agent through the relevant API

I integrated a number of different cloud platforms for this work, including Magnific, Beeble and Slapshot. In the end, most generative passes were done directly by a Cursor agent through the relevant API

Rough roto of the rider was done using Slapshot.ai. Plate is courtesy of user 'Corey Stowell' on Pexels.com, and is intentionally challenging - with a fast moving camera, lots of lens distortion and harsh lighing
Rough roto of the rider was done using Slapshot.ai. Plate is courtesy of user 'Corey Stowell' on Pexels.com, and is intentionally challenging - with a fast moving camera, lots of lens distortion and harsh lighing

Rough roto of the rider was done using Slapshot.ai. Plate is courtesy of user 'Corey Stowell' on Pexels.com, and is intentionally challenging - with a fast moving camera, lots of lens distortion and harsh lighing

Camera tracking was automated locally using a custom tool and agent skill. Tracking is ideal to run inside an agentic loop, since 'success' is largely objective, and the agent can fine tune parameters or elevate to a more sophisticated technique after each loop

Camera tracking was automated locally using a custom tool and agent skill. Tracking is ideal to run inside an agentic loop, since 'success' is largely objective, and the agent can fine tune parameters or elevate to a more sophisticated technique after each loop

The target environment is a desert wadi, surrounded by large sandstone rock formations. This is the template script for the first modular asset - the initial blocking is generated in Rodin through a Griptape script, detailed in Houdini and baked in Substance

The target environment is a desert wadi, surrounded by large sandstone rock formations. This is the template script for the first modular asset - the initial blocking is generated in Rodin through a Griptape script, detailed in Houdini and baked in Substance

Once we have a working template, we can duplicate and re-run it for an arbitrary number of inputs. In this case, the only manual step was specifying a set of reference images, and leaving the agent to run. The outputs are written as Unreal-compatible USD files

Once we have a working template, we can duplicate and re-run it for an arbitrary number of inputs. In this case, the only manual step was specifying a set of reference images, and leaving the agent to run. The outputs are written as Unreal-compatible USD files

A second canvas is used as a material library. Any node can be clicked on to view a small embedded 3D material viewer to see it applied in context. Another custom skill and Unreal bridge plugin is used to automatically set up a layered shader in unreal
A second canvas is used as a material library. Any node can be clicked on to view a small embedded 3D material viewer to see it applied in context. Another custom skill and Unreal bridge plugin is used to automatically set up a layered shader in unreal

A second canvas is used as a material library. Any node can be clicked on to view a small embedded 3D material viewer to see it applied in context. Another custom skill and Unreal bridge plugin is used to automatically set up a layered shader in unreal

The primary control surface is an Unreal Environment. The basic layout was done rapidly using the modular rock assets, with a secondary procedural erosion and scatter pass in Houdini (see the 'Greenhouse' page for more info on this)
The primary control surface is an Unreal Environment. The basic layout was done rapidly using the modular rock assets, with a secondary procedural erosion and scatter pass in Houdini (see the 'Greenhouse' page for more info on this)

The primary control surface is an Unreal Environment. The basic layout was done rapidly using the modular rock assets, with a secondary procedural erosion and scatter pass in Houdini (see the 'Greenhouse' page for more info on this)

Creating an environment like this, with pre-loaded assets and a material library, is very fast. In my testing, the time up front to create an actual 'world model' saved countless interations, prompting hacks and convoluted layout diagrams, but this will obviously vary by project

Creating an environment like this, with pre-loaded assets and a material library, is very fast. In my testing, the time up front to create an actual 'world model' saved countless interations, prompting hacks and convoluted layout diagrams, but this will obviously vary by project

The secondary control is from a single styleframe, demonstrating how the rider should be integrated into the scene. I experimented with painting this both with and without motionblur, and the 'no motionblur' was more successful
The secondary control is from a single styleframe, demonstrating how the rider should be integrated into the scene. I experimented with painting this both with and without motionblur, and the 'no motionblur' was more successful

The secondary control is from a single styleframe, demonstrating how the rider should be integrated into the scene. I experimented with painting this both with and without motionblur, and the 'no motionblur' was more successful

Cheap precomp for all subsequent generations. The lens distortion was so extreme I didn't render the full overscan to save time, and left the black corners to be filled in later

I tried all the available video-to-video, multi-modal or ref-to-video models currently available, with honestly quite disappointing results. The fast moving camera seems to cause the styleframe to be ignored, and prompt adherance was relatively poor. These are Kling O3, LTX+LoRA, Minimax H3 and Seedance 2.5

All generations were recorded on the canvas for tracking, including seed, ID and approximate cost

All generations were recorded on the canvas for tracking, including seed, ID and approximate cost

Beeble's SwitchX is very promising from a 'pixel perfect match' standpoint, but it also doesn't seem to handle vehicle interaction or fast moving cameras very well

Gemini Omni on it's own gave natural looking results, but would also occasionally throw very random outputs for no discernable reason. Running SwitchX first made the layout and style adherance more solid and predictable

I was really hoping to get compositing-friendly adherance to the layout, so different generations could be combined. This seems very possible with a still or slow moving camera, but for this shot, every model produced very slightly different compositions - even when given identical seeds

This was the Gemini Omni generation that was most successful, though I would have plenty of notes if this was a shot intended for delivery. It's also (like Seedance 2.5) currently limited to the relatively low resolution of 1280x720

Comparison between the Omni output and the pre-comp. Close, but not as close as I was hoping, and certainly not 'comp friendly'. I tried several different retake, conform and chaser generations, but no tool or model would align a natural-looking shot with decent integration back to the layout

Comparison between the original plate and the Omni output. Overall, there are some small glimpses of a solid, production-friendly workflow here - but not something I would be happy delivering as a final. I now have the workflow and the harness to run this at scale, so I'll be on standby for model updates that fit into this system

A second test was the opposite requirement - adding a creature into a plate, while retaining as much of the original detail and quality as possible

For this I had access to the location, so part of the test was conforming a photogrammetry scan to a camera track automatically. The same 'agent loop' tooling was used, with an additional iterative closest point (ICP) process to match the tracked point cloud with the mesh

For this I had access to the location, so part of the test was conforming a photogrammetry scan to a camera track automatically. The same 'agent loop' tooling was used, with an additional iterative closest point (ICP) process to match the tracked point cloud with the mesh

The scene reconstruction removed a lot of the scale ambiguity, and allowed for relatively accurate world space layout and blocking. Low-res raven asset is courtesy of FAB

The scene reconstruction removed a lot of the scale ambiguity, and allowed for relatively accurate world space layout and blocking. Low-res raven asset is courtesy of FAB

The previs was then skinned in NanoBanana, with some manual DMP work on top to bring back the details from the original plate, and improve the claw contact

The previs was then skinned in NanoBanana, with some manual DMP work on top to bring back the details from the original plate, and improve the claw contact

The same process was repeated for the last frame
The same process was repeated for the last frame

The same process was repeated for the last frame

Many of the leading models were disappointing, required dozens of iterations, and were very expensive. It wasn't that the quality was necessarily that bad, but adherance to the plate was often very poor. This is Kling O3 and Seedance 2

Again, slightly to my surprise, Gemini Omni proved to have the best balance of natural movement, integration, plate and prompt adherance. Limited to 720p initial resolution (before upscaling) as a result.

Slider comparison of the plate against the Gemini Omni generation, showing the close but not pixel-perfect match between the generation and the plate