By now most of you have probably heard about the launch of PLAI by GEMA, and that it landed eight days before the first-instance ruling in the lawsuit against Suno. That timing is clearly no coincidence. The German society wants to show that AI training and fair compensation for creators are not mutually exclusive.
So what exactly is this new product?
It's a licensed music catalog built for training AI models. It currently holds around 178,000 audio files across roughly 57,000 works, spanning more than 60 genres. Worth noting: this is production music, meaning tracks composed specifically for sync use in advertising, film, TV, and other audiovisual content.
What's new about this dataset is that it aims to let AI companies clear everything in one place. Training a model usually means clearing two separate layers of rights: the copyright in the composition and the master rights in the specific recording. PLAI bundles both, along with the audio files and metadata.
Beyond basic genre, tempo, and instrumentation data, it includes track structure, ISWC, ISRC, emotion tags, and use-case context, the kind of information a model needs to learn beyond the audio itself. GEMA builds custom training packages depending on the model and the client's use case.
That points to a real understanding of client needs, which matters if collective management organizations want to negotiate on equal footing with the tech giants.
The project is far from perfect, and I'll get into why in this article, but it's a first step in the right direction.
What kind of clients is PLAI targeting?
The criteria for who gets to buy in are explicit: AI tools that assist creators during the production process, and whose output doesn't compete with the works used to train them. It's a clear dividing line between two worlds. On one side sit generative models that compose full songs from a prompt (Suno and Udio's territory); on the other, tools that help a musician or producer work better without replacing their creative role.
Klangio is the first named client, and it fits that criteria well. Its tools turn recordings into sheet music. They don't generate new music, they read music that already exists.
What's the main criticism it's getting?
Matthias Strobel Hohmann, chairman of MusicTech Germany, was harsh on the catalog as soon as it launched.
"An AI model trained solely on stock and library music does not learn songwriting, emotions, or the stylistic diversity of popular music. 178,000 audio files is a joke."
He has a point when it comes to generative models, which need millions of tracks to produce acceptable output. But that's not what PLAI is for. Its whole premise is that its clients' tools shouldn't compete with the works used to train them.
Another open question is how creators will actually get paid. GEMA says compensation will be proportional, but the formula hasn't been published. My take is that it's still early, this is very recent and there's no established standard yet. Still, it's a critical issue for GEMA's members, and pressure will keep building if the silence drags on.
Is this the first licensed dataset for AI training?
One thing I looked into while writing this is how novel the project actually is. The answer is: it depends. Among collective management organizations, GEMA is unquestionably the first to do something like this. The German society had actually flagged two years ago that it was working on the issue of creator compensation in an AI context.
But step outside collective management and you find Rightsify and vAIsual, which have partnered since 2023 to sell licensed music catalogs built specifically for AI training. Rightsify supplies the catalog with rights cleared; vAIsual packages and distributes it through its Dataset Shop platform. So the concept of a licensed music dataset already existed before PLAI.
The difference is institutional origin. Rightsify is a B2B licensing company operating outside the collective management system, while PLAI comes from a CMO that pools its publisher members' repertoire under the same mechanism it uses for the rest of its operations.
What comes next?
GEMA now has to solve the chicken-and-egg problem every marketplace faces. It needs the catalog to grow to attract more clients, and it needs more clients to attract more rightsholders. On its own website, GEMA says it's flexible about bringing in more publishers, societies, authors, and other rightsholders.
One thing I keep coming back to is enforcement. How do you actually verify that a model was trained purely on licensed catalogs? Right now that's nearly impossible, like trying to stop rain with your hands. Attribution models exist, like the one built by Sureel AI, but I'm somewhat skeptical about how effective they really are. Nobody's addressed this yet, though I have no doubt the question will surface sooner or later.
Closing thoughts
As I write this, we're four days out from the Munich court's ruling on GEMA's lawsuit against Suno. Whatever happens on Friday will shape this project going forward. A ruling in GEMA's favor puts this on the fast track. If it goes the other way, I doubt PLAI disappears, but I'd expect its potential impact to shrink.
GEMA is clearly positioning itself as the global reference point in the fight against big tech, and it's paving the way for others to follow. I wouldn't be surprised if the rest of 2026 brings announcements from other societies laying the groundwork for their own licensed catalogs.
Sources: AI Musicpreneur; Velveteen; Musically; GEMA; Music Business Worldwide