Amazon Devices logo
Amazon Devices

Devices Metadata Studio

Built a series of AI experiments to test what metadata could be inferred from images, what couldn't, and where the tradeoff between automation and human input actually belonged.

Devices Metadata Studio

Role

Lead UX Designer

Timeline

August 2025 to Today

Team

PM, ML Engineers, Data Science, Compliance

Designer note

As we learned building Image Builder, every layer in a campaign image, the device, the background, the copy, needs metadata attached to it. Each image had 30 to 90 metadata inputs, many of them manually entered, which was the tradeoff we'd accepted to keep that knowledge in-house.

Going against that predicament, I always thought there's gotta be a better way, so rebelliously, I built a series of small experimental tools to test what AI could figure out on its own from an image, what it couldn't, and where it was reliable enough to actually trust. Ultimately got buy-in from stakeholders, and these AI-powered processes are now part of the workflow, without a design technologist manually entering data.

The Problem

Getting the metadata wrong could mean automated errors with truly severe consequences, a US power plug in a German ad, effectively breaking the law, a dog in a UAE campaign (dogs aren't loved there), licensed movie content on a screen in a market where it wasn't cleared. This doesn't happen with the set of eyes that go through every image manually created for all these markets.

The work was manual, the team was small, and the data was proprietary, so we couldn't outsource it to an off-the-shelf AI tool. This manual entry had become the single biggest bottleneck in the whole asset pipeline.

The manual upload screen for a background image. Up to 90 metadata inputs, before a single asset was ready.

Design and Build

To focus engineering efforts, I mapped the various component image types against their metadata requirements and potential for AI automation. This matrix became my framework for deciding where AI could replace manual metadata work, where it couldn't, and where engineering effort was worth investing.

My framework for prioritizing where AI could replace manual metadata work, and where it couldn't.

With the matrix identifying the highest-value opportunities, I tested specific possibilities through small experiments and prototypes. Not everything worked. A mini-tool I built to blindly classify raw component images failed. But another prototype successfully matched component images to predefined visual variants, proving the concept was viable. These weren't abstract explorations, they were concrete, hands-on tools built to make the argument with evidence before asking Engineering to invest in production functionality. This was early 2025, when AI capabilities were significantly less mature than they are now.

A failed experiment. Tried to get AI to classify raw component images with no existing metadata to lean on.

This one worked. AI matched component images to predefined visual variants.

The experiments established what was technically feasible, and the resulting product direction reduced the metadata users needed to provide manually. Critical metadata is shown first. Deeper metadata is available through progressive disclosure when needed. The user is not overwhelmed by 30–90 fields at once. AI is used where it can reliably infer information, and human input remains where it is still necessary.

A screen mapper tool was introduced to give the AI the spatial information needed to place screen content correctly on physical devices. Users define the digital boundary on a physical device, and the system translates that into coordinate data for dynamic composition.

Users drew the boundary on the device. System turned that into coordinates the AI could actually use.

Concept for the Admin Tool. The question was simple: how much metadata could AI handle so people didn't have to.

Result

The experiments proved AI could handle enough of the metadata burden to change the product direction. Product, Engineering, Data Science, Compliance, and Brand aligned around a shared roadmap. Three capabilities greenlit: Visual Variant Attribution, Devices Screen Mapper, and AI-First Screen Image Uploader.

The screen mapper and AI composition work proved localized screen content can be placed into device imagery correctly. The remaining work, the UI for mechanized metadata ingestion, is in build.

Lord of the Rings, placed inside an Echo Show. Same process localizes screen content, keeps perspective, shadows, and glare intact.