Cloud
Engineering
September 28, 2026

10,195 AI clips in one weekend, powered by clean energy

Monks, a global marketing and technology services company, set out to give Boomtown attendees a video with their own name on it. They built two open source pipelines on NVIDIA H200 GPUs in Crusoe Cloud's Iceland region and shipped 10,195 5-second clips in one production window. Here is how.

Rohit Kalmankar headshot
Rohit Kalmankar
Senior Staff Solutions Engineer
September 28, 2026
Isometric illustration of server racks and a metrics dashboard representing NVIDIA HGX H200 systems on Crusoe Cloud

Boomtown is not a typical music festival. Known for its intricate world-building, sustainability commitments, and immersive storytelling, it draws tens of thousands of attendees every year to a temporary city that exists for just a few days.

In 2026, Monks, a global creative technology company, set out to match that ambition with something new: sharing an aftermovie for every visitor (~77,000 in total) within 48 hours after the festival, providing a memory about everyone's individual festival journey. For a specific part of the videos they explored new ways to create personalized content, powered by NVIDIA HGX™ H200 systems connected with NVIDIA Quantum-2 InfiniBand networking on Crusoe Cloud in Iceland.

The result: 10,195 uniquely personalized flag clips and thousands of AI-generated cinematic transitions, produced in days, not weeks, with a fraction of the manual effort traditionally required.

This is how they did it.

Solution architecture

Isometric architecture diagram of the Boomtown generative video pipeline on NVIDIA HGX H200 systems in Crusoe Cloud

The challenge: video production at festival scale

Producing personalized video content for a live event has always been slow, expensive, and impossible to scale. The traditional approach to 10,000 personalized five-second clips would look something like this:

  • Human effort: 3 to 5 full-time creatives
  • Commercial value: >$30,000
  • Tooling: Adobe After Effects and licensed creative software
  • Turnaround: 3+ weeks

For a festival that lasts a weekend, that timeline simply doesn't work. And personalization, giving every single attendee their own video, was effectively impossible at any reasonable cost.

Monks believed there was a better way. But to prove it, they needed serious GPU compute to experiment with heavyweight open-source AI models that simply couldn't run on local hardware.

Why Crusoe, and why Iceland

Before a single video was generated, Monks needed an infrastructure partner that could give them something other cloud providers couldn't: access to enough NVIDIA HGX H200 nodes to run large-scale model R&D and production workloads, without the provisioning barriers they had experienced elsewhere.

From the start of the engagement, Crusoe's solutions engineering team recommended the NVIDIA HGX H200 in Crusoe Cloud's Iceland region as the right infrastructure for this workload, before the pipeline was even designed. That recommendation was based on three factors: the NVIDIA H200 GPU's 141GB HBM3e memory making it ideal for running large open-source generative models at full precision; Iceland's geographic proximity to the UK festival for faster data transfer; and Crusoe's ability to provision dedicated, isolated GPU nodes at scale without the queue times the team had faced on other platforms.

"Crusoe gave us access to more GPUs, we didn't have the same barrier we had trying to provision the ones we needed for this R&D." - Monks team

That early infrastructure decision turned out to be the right one. The NVIDIA H200 GPU nodes handled sustained, mixed image-and-video generation for days without a single hardware-related failure.

The choice of Iceland was equally deliberate beyond latency. Boomtown is a sustainability-first festival. Its audiences care deeply about environmental impact, and AI-generated content powered by fossil fuels would have been at odds with everything the festival stands for. Crusoe's Iceland infrastructure runs on clean, renewable energy from geothermal and hydroelectric power.

NVIDIA's GPU-accelerated compute stack, specifically the H200 with its 141GB HBM3e memory and NVIDIA Quantum-2 InfiniBand networking, provided the raw power needed to run the open-source models at the heart of both pipelines.

The infrastructure

Crusoe provisioned two dedicated NVIDIA HGX H200 instances in eu-iceland1-a, a private VPC with no shared tenancy, reserved specifically for this production run.

  • Cluster: 2x Crusoe instances, 8x NVIDIA H200 GPUs each (16 GPUs total)
  • GPU memory: 141GB HBM3e per GPU, 1,128GB per instance, 2,256GB total across the cluster
  • CPU / RAM: 176 vCPUs (Intel Xeon Sapphire Rapids), 2,000 GB RAM per instance
  • Region: eu-iceland1-a, geothermal and hydroelectric power, closest Crusoe region to the UK
  • Networking: NVIDIA Quantum-2 InfiniBand, 3,200 Gbps, <1ms intra-cluster latency, 175 Gbps VPC
  • Storage: 8x 1.92TB NVMe per instance
  • Data transfer: ~398 MB/s parallel inbound (Iceland to Frankfurt)
  • Models: Open source, no proprietary software licenses, no vendor lock-in
  • Orchestration: Resume-safe job queues with deterministic sharding across both nodes

The NVIDIA Quantum-2 InfiniBand networking between instances meant both nodes could operate as a single coordinated cluster, which mattered for splitting the 10,195-name workload across parallel GPU queues with no duplication or data loss on any node failure.

Pipeline A: the overgrowth clip-transition system

The first challenge was creative continuity. A festival film is built from hundreds of clips shot across multiple stages, days, and moments. Cutting between them with a hard cut or a simple dissolve loses the magic. The Monks team wanted something more, a transition that felt alive.

The solution: grow nature between clips.

Between any two festival clips, the AI generates a short "overgrowth" bridge of ivy, mushrooms, flowers and butterflies that morphs one clip into the next and then hard-cuts back into real footage. The whole transition takes 8 to 9 seconds and is entirely AI-generated from real frames, running as a batch job across the NVIDIA HGX H200 nodes on Crusoe Cloud.

How it works

Rather than describing the transition in words, the model is given four real frames as anchors: two from the end of clip A and two from the start of clip B. The model treats the A pair as "before" and the B pair as "after," then generates every frame in between as one continuous sequence.

The key insight: two frames per side, not one. A single frozen frame per side causes the crowd and camera to slow to a standstill mid-transition, so it looks like the video hits pause. Two frames a fraction of a second apart carry velocity information, so the crowd keeps moving naturally through the overgrowth.

To ensure the generated transition matches the real camera movement, hundreds of tracking points are mapped across each clip. When almost all trails sweep the same direction, that shared motion is the camera move. This tells the model which frames to use as guides.

At the landing point, the AI tends to drift. Rather than trusting the model's last frame, the pipeline automatically finds the generated frame that pixel-matches real clip B, and cuts there. No dissolve and no crossfade. The motion simply continues.

Three rotating motifs keep a long sequence of transitions feeling fresh: ivy and moss, orange mushrooms, and flowers with butterflies. It all runs as a batch job on the NVIDIA H200 GPU nodes, thousands of clips at a time.

Pipeline B: personalized flag videos at scale

The R&D phase, testing open-source models on Crusoe's NVIDIA H200 GPU nodes and identifying which performed best at each task, informed the design of the scaled production pipeline. The hypothesis: mass-produce thousands of uniquely personalized flag videos, each carrying an individual attendee's name rendered onto real flag fabric, using a two-stage pipeline running in parallel across 16x NVIDIA H200 GPUs.

The models

  • Qwen-Image-Edit, renders each attendee's name onto flag fabric respecting folds, lighting, and perspective
  • LTX-2.3, animates the static flag plate into a 5-second, 121-frame waving video

Both are open-source models downloaded directly onto the Crusoe instances. No licensing fees. No API calls. The NVIDIA H200 GPU's 141GB HBM3e memory allowed both models to run at full precision, ensuring consistent quality across all 10,195 unique name renders.

The pipeline

Stage 1, image generation (Qwen-Image-Edit)

Qwen takes a photographic flag plate and renders the recipient's name directly onto the fabric, respecting the folds, lighting, and perspective of the cloth. The result looks like the name was printed there.

Stage 2, video generation (LTX-2.3)

LTX-2.3 animates the static flag plate into a 5-second, 121-frame video of the flag waving in the wind. A final lightweight crop stage packages the deliverable.

Everything runs across three color variants, red, green, and purple, with fixed seeds per stage so the same name always reruns to the same pixels, with no sampling artifacts at scale.

Running it on Crusoe

The name list was split across the two 8x NVIDIA H200 GPU instances by deterministic sharding, so any node could crash and resume without duplicating or dropping work. Every GPU pulled jobs from a resume-safe queue and wrote results to a per-node ledger, giving live visibility per stage: completed, retries, pending, success rate, and cumulative compute time.

The team's approach: batch testing first, always. 10 images, then 50, then 100, then a quality check, then 10,000. Then repeat for video. This phased approach, made possible by having dedicated Crusoe nodes available for testing before the production window, meant the live run was smooth and predictable.

A text recognition (OCR) script ran on the instances after image generation to verify every name was correctly rendered before the video stage. Post-production cropping ran as a script directly on the Crusoe instances, eliminating a separate post-processing step entirely.

Final delivery: completed videos were pulled directly from the Crusoe instances into the broader client delivery pipeline, matched by filename to each attendee.

By the numbers

MetricResult
Unique names in final production10,195 (red 3,776 / green 3,210 / purple 3,209)
Total videos generated10,000+ including test batches and seed sweeps
Success rate100% on standard names, zero abandoned jobs across all stages
Median image generation time~33 seconds per personalized image (NVIDIA H200 GPU)
Median video generation time~54 seconds per 121-frame video (NVIDIA H200 GPU)
Median crop time~1.8s per crop
Total compute~226 GPU-hours across the cluster
Cluster2× Crusoe instances · 8× NVIDIA H200 GPU each · 16 GPUs in parallel

Before vs. after

Traditional (no AI)AI on Crusoe’s NVIDIA H200 GPU
People required3–5 FTE1–2 FTE
Commercial value>$30,000~$20,000
Turnaround3+ weeks1–2 weeks
ToolingAdobe After Effects (licensed)Open source (Qwen-Image-Edit + LTX-2.3)
PersonalizationNot feasible at scale10,195 unique videos
Cost per video~$2.94~$1.96

At first glance the cost difference appears modest. But the economics compound quickly at scale. At roughly 33% savings per video (~$2.94 traditional vs ~$1.96 on Crusoe), running the same pipeline across multiple festivals per year translates to over $100,000 in annual savings, while unlocking personalization at a scale that was simply not achievable with the traditional approach.

Importantly, the cluster had headroom remaining after the production run. That means the same infrastructure could have processed significantly more videos within the same window, further driving down the cost per video. The pipeline setup cost is largely fixed. The more volume you push through it, the better the unit economics get.

What failed and what we learned

A small subset of names, under 1%, mostly accented characters and unique spellings, needed image plates from an alternative image generation tool. Those plates froze in the video stage, producing a static flag instead of motion. Root cause: plate compatibility between image sources and the video model, not a flaw in the pipeline or the hardware.

The NVIDIA H200 GPU cluster ran sustained, mixed image-and-video generation for days without a single hardware-related failure. Every incident debugged during the production run was on the software side. With the infrastructure behaving predictably, the team could focus on the creative and software problems.

What the two production runs proved together

Both pipelines scale the same way: add Crusoe nodes, reshard the queue. The tooling built here is a reusable template for future generative video products, for Boomtown 2027 and beyond, and for any live event that wants to give its audience something personal to take home.

"The nodes carried a heavy generation load exactly when we needed them, and that made a real difference to the project." - Monks team, Boomtown 2026

The 10,195 personalized flag videos were one part of a much larger Boomtown production. Across both pipelines, Monks rendered approximately 180,000 clips in total, delivering 77,000 videos to attendees, 10,195 of which were uniquely personalized with each individual's name.

The infrastructure recommendation that started this engagement, H200s in Iceland, made before a single line of pipeline code was written, proved to be the right foundation for everything that followed.

Open source. Clean energy. Zero hardware failures.

No proprietary software licenses. No vendor lock-in. Two open-source models (Qwen-Image-Edit and LTX-2.3), 16 NVIDIA H200 GPUs on Crusoe's renewably-powered Iceland infrastructure.

That's what it took to give 10,195 festival-goers something they'll never forget.

Conclusion

Boomtown 2026 proved that AI-generated, personalized video content at festival scale is no longer a concept. It's a repeatable, cost-effective production pipeline. Built on open-source models and NVIDIA HGX H200 systems in Crusoe Cloud's Iceland region, the pipeline developed here is ready to go again, for the next festival and the next creative brief.

Thanks to the Monks technical team for their work on the pipelines.

Building a generative video pipeline of your own? Explore Crusoe Cloud, or talk to our team.

Latest articles

Chase Lochmiller - Co-founder, CEO
September 28, 2026
10,195 AI clips in one weekend, powered by clean energy
Chase Lochmiller - Co-founder, CEO
September 21, 2026
Crusoe earns NVIDIA Exemplar Cloud validation on Blackwell Ultra
Chase Lochmiller - Co-founder, CEO
September 16, 2026
Serving 5.75 million tokens per second: Crusoe's MLPerf Inference v6.1 results on AMD MI355X

Are you ready to build something amazing?