Connect with us

Net Influencer

Tech

How Diffraction Is Building Creator Rights Into Physical AI’s Training Data Market

AI labs and robotics companies need training footage to teach machines how to grip tools, navigate uneven terrain, or judge how spilled milk moves differently from water. Much of that footage is being filmed daily by content creators who neither know it has commercial value nor are being compensated for it. Diffraction, the Lancaster, Ohio -based Influencer Marketing and data company, is building the contracting, capture, and provenance infrastructure to change that arrangement.

Joseph Sottile co-founded Diffraction in September 2022 and launched StreamGenie, the company’s data layer, in January 2026. The company runs influencer campaigns and a TikTok LIVE creator network in the U.S. and UK, producing an estimated 40,000 to 50,000 hours of live content per month. “We noticed that creators were sitting on archives that no one was paying for, but were being used for training already,” Joseph says. “We wanted to connect the dots.”

Co-founder Thom Swann is behind the technical infrastructure to act on that observation. He previously built Twitch extensions still in active use today and engineered Wendy’s digital ordering and kiosk program. His view of the physical AI opportunity sharpened when Diffraction produced its first purpose-built dataset, built to spec for a specific research lab rather than pulled from public archives.

The company’s three business lines run on a single underlying asset: relationships with people who make video. “Marketing monetizes a creator’s reach,” Joseph says. “Licensing monetizes their work and the videos they’ve created. Capture programs monetize their time.” Custom data capture for physical AI and robotics, the most recently developed of the three, is where the company has seen the sharpest increase in inbound demand.

Rights and Consent Are the Core Product

Joseph and Thom believe the physical AI data market has a fraud problem. Footage is screen-recorded, lifted from public sources, or aggregated without clear rights documentation, practices that undermine the legal integrity of any dataset built on them. “The industry has a real big fraud issue right now with people screen recording, lifting videos, doing things that just aren’t giving the best source of light right now,” Thom says.

For Diffraction, provenance is the first design constraint. Every dataset the company builds or licenses today carries a documented rights trail showing that the content owner consented and granted the specific rights being used. That discipline has recently become a legal requirement in one major market. The EU AI Act has required providers of general-purpose AI models since August 2025 to publish a summary of the content used to train their models. “To publish that,” Joseph explains, “whatever you’re training on, you need to be able to legally and contractually show that the rights were in fact given.”

Diffraction aims to build the same protections into every license it issues. Use of likeness is prohibited. So is one-to-one replication, and duplication of the licensed data. And the terms require deletion once the licensed use is complete, so a one-time purchase does not quietly become a permanent asset. “You don’t want to license your content, give someone the rights, and then all of a sudden they’ve duplicated it, and they don’t need you anymore,” Joseph says. 

The obligation extends to third-party elements embedded in creator footage, including copyrighted music playing in a stream’s background or other subjects who appear on camera and must separately consent.

Diffraction Builds Datasets to a Buyer, Not a Catalog

A large portion of the physical AI data market operates on a speculation model: aggregate broad volumes of footage across general categories and wait for a buyer. Diffraction works in the opposite direction. “Rather than bucketing a bunch of it and hoping that it gets sold, we build it to spec,” Joseph says. “So it is sold, because we have a buyer for it.”

How Diffraction Is Building Creator Rights Into Physical AI’s Training Data Market

The company has conducted more than 500 content audits, developing a working map of what creator content draws active buyer demand. A creator with 100 gigabytes of archived streams might find that 20 hours of on-camera LEGO building is the portion with immediate market value, while the rest sits outside current requirements.

Pricing ranges from single digits to several hundred dollars per hour, a gap explained by three variables: scarcity, skill, and variation. According to Joseph, footage of everyday tasks, available in abundance across the internet, carries the lowest value. A licensed plumber performing an active repair commands a premium, because the model learns not only from the correct outcome, but from the contrast between a skilled practitioner’s efficient path and an amateur’s inefficient one. Depth capture, annotation layers, and footage recorded across multiple environments add further value.

Thom notes that bulk aggregation also carries an environmental cost. Storing footage speculatively consumes compute and energy, waste that compounds when the capture spec is not yet refined enough to know what buyers actually require. Building against a known requirement sidesteps that cost on both sides.

How Diffraction Is Building Creator Rights Into Physical AI’s Training Data Market

Physical AI Training Data Is Not About Who You Are

The most persistent misconception Joseph encounters when introducing creators to this market is that licensing footage for AI training means surrendering their likeness. The fear is that a research lab will build a digital replica and deploy it without further compensation. “Couldn’t be further from the truth,” he says.

What physical AI and robotics buyers purchase, in his words, is behavioral and physics data: how a hand grips a tool, how light reads across different surface textures, how a robot should respond when the environment changes unexpectedly. The person on camera, their face, name, and appearance, is irrelevant to the dataset’s value. “Your likeness, who you are, what you look like, everything outside of what you’re actually doing, doesn’t matter,” Joseph explains. 

Deals Diffraction has been a party to, including footage purchases in the range of millions of hours, treat content as training substrate. Who appears in it is not part of the transaction.

Diffraction’s licenses prohibit likeness use and replication, and require deletion after the licensed use. Joseph expects a coming wave of litigation and enforcement to sharpen the distinction between compliant licensing and unauthorized use. “There’s no accountability for AI right now,” he says. “Some of the first lawsuits and court hearings are coming through, and that’s going to be a lot over the next year or two.”

How Diffraction Is Building Creator Rights Into Physical AI’s Training Data Market

From Filming to Operating: The Teleoperation Roadmap

Diffraction’s roadmap runs through three stages. The first, licensing archival footage, draws on what already exists. The second, custom capture to spec, builds what buyers require. The third, teleoperation, places operators remotely in control of robot arms to generate demonstration data, direct recordings of a human performing a task through a robot’s mechanism, which is among the highest-value inputs for physical AI development.

“It starts with licensing what exists on the internet,” Thom says. “Moving into capturing what buyers actually need via spec, and then ending at teleoperation where people are remotely operating the robot arms.”

Getting there requires two things: a workforce comfortable with VR control interfaces and research lab partners willing to run pilot programs. Teleoperation rigs typically use VR headsets, and the skill of navigating a virtual environment transfers directly to operating a robot arm remotely. “There are a lot of people in the United States that have bought [Meta] Quests for their kids,” Joseph says, “and can start helping shape, via their VR Quest headset, some form of teleoperation.”

The arc is also a qualification path. Thom says a creator who films their own work to spec today is learning the things that it takes to teleop tomorrow.

The Creator Who Builds Around a Skill Has the Most to Gain

Joseph’s near-term prediction centers on robotics, brand marketing, and creator content. He points to a robot bartender on a Royal Caribbean cruise he encountered five years ago that could initially make five drinks and has since learned to take custom orders through continued training. The brands deploying such technology, he expects, will want creator involvement, both to drive audience engagement with the experience and to generate interaction data simultaneously.

That creates a specific advantage for a specific kind of creator. “Pick something that you’re good at, that you’re passionate about, and actually have a skill in, and build content around it,” Joseph says. “Because you’re skilled, you’ll have more credentials to use that in this next age of training data.”

Subscribe to Our Newsletter


Check Out Our Podcast

Tamara Blazquez

Tamara is a writer, editor, and project manager passionate about using storytelling to inspire awareness, connection, and positive change. With years of experience leading creative teams, developing global campaigns, and producing award-winning visual and written stories. As Impact Storytelling Manager at Photographers Without Borders, Tamara managed an international team of writers, designers, and photographers, coordinating content creation, editing, workshops, and grant programs focused on social and environmental impact. Her work as a freelance travel writer for Static Media's Islands further sharpened her research and editorial skills while deepening her understanding of global tourism, culture, and sustainability.

Click to comment

More in Tech

Tips, Comments, Suggestions? Email Us!

[email protected]
To Top