At Higgsfield AI, we’re building generative video products tailored for creative professionals — and high-performance, efficient infrastructure is foundational to what we do. As part of that mission, we developed Higgsfield DoP I2V-01-preview, our proprietary Image-to-Video (I2V) model designed for high-quality video generation from images. In collaboration with TensorWave, a cloud provider offering AMD-based compute instances, we benchmarked and profiled our model on AMD Instinct™ MI300X GPUs, evaluating its performance across inference workloads and real-world deployment scenarios. Higgsfield DoP I2V-01-preview brings cinematic structure and control into generative video through a novel architecture that blends diffusion models with reinforcement learning. Rather than simply denoising frames, the model is trained to understand and direct motion, lighting, lensing, and spatial composition — capturing the grammar of cinematography. Inspired by how reinforcement learning has been used to teach large language models reasoning skills, we applied RL after diffusion to instill intent and coherence in generated sequences. The result is a system capable of producing expressive, controllable, and high-fidelity video — powered by a robust infrastructure stack purpose-built for professional creative workflows.
Inference with Tensorwave on AMD Instinct™ GPUs
TensorWave’s AMD-based infrastructure, built on next-generation accelerators, provided a scalable and memory-optimized environment ideally suited for our inference workloads. With pre-configured PyTorch and ROCm™environments, we were able to run inference out of the box — no custom setup or engineering effort required. This streamlined deployment experience allowed us to focus entirely on validating the stability and performance of our model on AMD Instinct™ MI300X GPUs.
Profiling I2V Inference: Where the Time Goes




