Back to Insights
6 min read

Attribution models: how they apply to copyright

Also available in:

In the middle of all the excitement around the ruling in GEMA's favor against Suno, I want to take a pause and talk about something most people have probably heard of but likely don't fully understand: attribution models.

My goal with this article is to explain, in simple terms, what they are, what they're for, and where this technology actually stands today.

An attribution model is, at its core, a technology designed to analyze a file generated with artificial intelligence and try to determine which works may have influenced it. We'll build on that definition as we go.

That file could be a song, an image, or a video. The whole point of attribution models is to work backward and figure out which training works shaped that specific output.

I'm genuinely uncertain about how reliable this technology can get. But if I've learned anything in over a decade working in tech, it's that nothing is impossible. Still, I'll walk through some technical and conceptual details later that explain why this is an extremely hard problem.

But first, what could you actually do with an attribution model?

One of the earliest ideas floated when people started discussing AI's impact was analyzing every single model output. That way, you could measure how much each work influenced the result and distribute royalties to creators proportionally. Over time, though, that idea lost steam because of how difficult it is to actually pull off.

It wouldn't just require an extremely high level of reliability. I genuinely can't picture a scenario where every response generated by every model in the world gets audited by an attribution model.

We're talking about billions of responses a day. Even if the analysis were extremely accurate, the computational cost would be enormous. And even if you assumed only a sample got audited, you'd still face a challenge just as hard, or harder: getting the whole industry to agree on a standard and getting companies to actually adopt it.

My sense is that there are better ways to approach this problem.

On the other hand, attribution models could work well as enforcement tools.

Take PLAI by GEMA, for example. An attribution model could analyze a generated output, identify which works influenced it, and check whether those works belong to the properly licensed catalog. Unlike an audio fingerprinting system, which only works if the output is a recognizable copy of the original, here the output is a new synthesis, so you need attribution to reconstruct that chain. In that context, an attribution model could become a tool for verifying license compliance, or even for producing evidence in litigation over unauthorized use of works.

Personally, this use case strikes me as far more reasonable.

Where does this technology actually stand?

The concept is still far from mature, but several companies are already trying to solve this problem.

Musical AI

Musical AI is a rights management platform for AI training and generation. On one hand, it lets rightsholders license their catalog under controlled terms, monitoring or restricting how it's used. On the other hand, it analyzes generated outputs and calculates how much of each result comes from which source, separating composition from recording.

In July 2026, SOCAN signed on as its first collective management organization partner. The deal covers two separate tools: consent, letting songwriters decide whether their work enters training, and attribution, measuring its influence on a specific output while distinguishing between composition and recording. It's the first time Musical AI has moved beyond one-off deals with labels and distributors into the collective management space.

Sureel AI

Sureel AI builds an "AI DNA" for each work, breaking it down into its components to track how AI models use it.

STIM, the Swedish music rights society, was the first to bet on this startup. In September 2025, it named Sureel its attribution provider, becoming the first collective management organization to trust its repertoire to this kind of technology. That set up the world's first collective AI license for music.

Warner bought Sureel this past June, aiming to own the technology that can prove how a catalog is being used. Sureel keeps operating as an independent platform, and in July, IMPF and IMPEL, two independent publisher organizations, announced they'll test it with a limited repertoire in a closed environment, with no use for training, before negotiating real terms.

ProRata

ProRata launched in 2024, founded by Bill Gross, the same person who invented the pay-per-click model for online advertising. It raised $25M in its initial round and another $40M in September 2025 led by Touring Capital. Since launch, it's had a strategic deal with Universal Music to explore applying its technology to music, but that deal hasn't turned into an actual product yet.

Where it does work is text. Gist Answers, its AI answer engine, splits 50% of revenue among more than a thousand publishers based on how much each source weighed into each response, and according to Gross it's already profitable. If they manage to bring that mechanism over to music, it would be the model closest to what most people picture when they think of attribution. For now, that's still the unfinished part of the business.

Pippa

Finally, and most recently, there's Pippa, a generative video startup launched in May 2026 that pays artists every time its tool generates content based on a style associated with them, $0.005 per image and $0.003 per second of video. It also sets aside a 5% pool from subscription revenue that participating artists draw from over time.

It's a much earlier-stage project than the others mentioned here: currently reports around 800 paying subscribers and just four artists with a signed deal, with four more in talks.

The CISAC president's take

Björn Ulvaeus, musician and president of CISAC, gave the opening address at the AI for Good Summit in Geneva, "The Future Needs Creators". If you haven't watched it, I'd recommend it. The message is powerful.

Björn doesn't seem to be a big fan of attribution models. He actually says: "For me, tracking the output was always the wrong question." My read is that he means they shouldn't be the primary mechanism for calculating distributions. His argument rests on how these models actually work: probabilistically. They don't produce copies of any song, they create new syntheses out of everything they've learned.

His proposal leans more toward catalog licensing and distributing revenue from AI subscriptions.

Closing thoughts

While I remain skeptical, attribution models could eventually reach a meaningful level of reliability and efficiency. Still, I don't think they'll play the starring role we imagined early on. More of an enforcement one.

The more I think about it, the more impractical it seems for output-based compensation to become the standard. My take is that the CISAC president, in his speech, was trying to steer the conversation away from that idea and toward what now looks like the more viable path: catalog licensing.

Related to
The Labs