Amazon Devices logo
Amazon Devices

AI Lifestyle Compositor

Built the AI compositor and human-in-the-loop review system that helped us generate campaign-ready lifestyle imagery at scale, using Creative Director judgment to train the engine and build trust in what it produced.

AI Lifestyle CompositorAI Lifestyle Compositor

Role

Lead UX Designer

Timeline

December 2024 to March 2025

Team

1 PM, 2 ML Engineers, 1 Creative Director, 2 Engineers

Designer note

We built and trained our own GenAI engine for lifestyle imagery, and I was responsible for designing the human-in-the-loop system that brought Creative Director judgment into the process. I worked with Brand Studio, Engineering, and Data Science on the feedback taxonomy, sampling model, and review experience, turning expert brand feedback into training signal for the engine.

This was early 2025. If I were approaching it today, I'd use agents to manage sampling, bring review directly to Creative Directors, interpret their feedback, and feed it back into the training loop. I'd probably only build dedicated UI for the moments where visual comparison actually needed it.

Review and approval experience.

The Problem

Lifestyle imagery converted about 70% better than gradient backgrounds in our internal data, but it was also one of the most expensive image types to produce. It depended on 3D artists, photographers, editors, and Creative Directors.

GenAI gave us a way to produce that imagery at much greater scale.

Amazon devices are confidential hardware, so external AI tools were off the table. We first experimented with an internal image-generation tool the Amazon Home team had built for placing furniture in home environments.

The results made the gap pretty obvious. Devices showed up at the wrong scale, with strange perspective, mismatched shadows, incorrect materials, or sitting in places where the product simply didn't belong.

We needed an engine that understood our devices and a way to teach it what campaign-ready creative looked like.

Lifestyle imagery converted about 70% better than gradient backgrounds in our internal data, but it was much more expensive to produce.

Lifestyle imagery converted about 70% better than gradient backgrounds in our internal data, but it was much more expensive to produce.

The furniture tool placed an Echo Pop in a room. Wrong perspective, flat device, mismatched shadows. The model simply did not understand our hardware.

The furniture tool placed an Echo Pop in a room. Wrong perspective, flat device, mismatched shadows. The model simply did not understand our hardware.

Building the System

We built the compositor on Amazon Bedrock using Amazon's proprietary device data.

I worked with Brand Studio, Engineering, and Data Science to define the product knowledge the engine needed: device families, camera angles, materials, proportions, shadows, screen reflections, and other rules specific to each device.

The system relied on two kinds of human input.

Design Technologists prepared background scenes and added structured information through Metadata Studio.

Creative Directors reviewed generated images and gave us the brand judgment the engine needed to learn from.

I focused on designing that second loop.

The engine could generate thousands of variations at once, so we needed a way to sample the output, capture Creative Director decisions consistently, and feed those decisions back into the model.

Every layer carries structured metadata. The device knows its product family. The background knows its environment type. The screen knows its campaign.

Every layer carries structured metadata. The device knows its product family. The background knows its environment type. The screen knows its campaign.

  • Brand-mandated camera angles, shadow rules, and copy space documented per device family. Turned into training data.

    Brand-mandated camera angles, shadow rules, and copy space documented per device family. Turned into training data.

    Echo Show in a defined scene with copy space. The model needed to know where text belonged before placing it.

    Echo Show in a defined scene with copy space. The model needed to know where text belonged before placing it.

Human in the Loop

I designed the review workflow, sampling model, and feedback taxonomy with Creative Directors and Data Science.

Some of the interactions were unusual.

A Creative Director could reject an entire batch and still need to identify the strongest image within that batch. That preference was useful training data, even though it wasn't a particularly natural thing to ask someone during review.

Early explorations used color-coded states, which became ambiguous quickly. We moved toward explicit rejection reasons and simplified the taxonomy into fewer, more useful options.

The review experience had to stay fast for Creative Directors while still capturing feedback detailed enough to improve the engine.

Each round of review gave the model another set of expert judgments to learn from, and the amount requiring human review decreased over time.

  • Went wide and unconventional trying to design something useful for science and non-intrusive for design directors. Ended up ambiguous, not accessible, and violated Meridian compliance. Taught me what not to build.

    Went wide and unconventional trying to design something useful for science and non-intrusive for design directors. Ended up ambiguous, not accessible, and violated Meridian compliance. Taught me what not to build.

    Shifted to explicit labels and reason-based feedback. Every state readable at a glance, Meridian compliant, no decoding required.

    Shifted to explicit labels and reason-based feedback. Every state readable at a glance, Meridian compliant, no decoding required.

Design to Code

This was also the first project where I shipped product code directly with Engineering.

I had started experimenting with Figma-to-code through MCP in August 2025, feeding design files into an AI-assisted workflow that generated working React. On Lifestyle Compositor, that experiment became part of the actual product-development process.

I built the batch review experience alongside Engineering instead of stopping at static design files.

There were terminal crashes and some angry GitHub comments along the way, but it worked.

It changed how I worked with engineers from that point forward. I later turned the process into a hands-on workshop for 14 designers.

Shared the MCP-powered Figma-to-code workflow with 14 designers in a hands-on workshop after shipping the review interface.

Result

The model kept getting sharper at understanding what a Creative Director would approve, so the amount requiring human review dropped with each round of feedback.

3,000Director-approved lifestyle assets shipped into Amazon Devices' image catalog*
10 hrs → 1.25 hrsCreative Director review time for batches of 12,000 AI-generated lifestyle images*

*Each review session presented four AI-generated variations of a prompt. The first pass sampled 5% of a 12,000-image batch, which meant 150 review sessions and about 10 hours of Creative Director time. As the model learned from tagged approvals and rejections, the required sample decreased. By the fourth pass, Creative Directors reviewed 19 sessions in about 1.25 hours, an 87% reduction in review time for the same output volume.

Sample of a real AI lifestly genearted image launched for a Fire Tv prmo in Sep 2025.