AI licensing raised the stakes
Media companies are now being asked to license their archives for AI use, and publishers, music labels, and studios are already signing deals. Each one comes down to the same four questions: what do we own, for which uses, in which territories, and until when?
Companies that can answer quickly can say yes with confidence and get paid. Companies that can't either walk away from revenue or sign terms they can't honor.
Regulation points at metadata
The rules are moving the same way. Under EU copyright rules on text and data mining, rightsholders must use machine-readable means, including metadata and the terms of a website or service, to reserve their rights. If your rights position lives in a PDF, it isn't machine-readable, and it isn't protecting you.
Embedding information in files isn't enough either. C2PA provenance metadata can be stripped by non-compliant platforms or by attackers. Rights data needs a trusted system of record, not just a tag on the file.
Small errors scale badly
Metadata errors don't stay small. As one music-rights analysis puts it, a single metadata error can be multiplied across thousands or millions of downstream uses, and once a model is trained, correcting it is extremely difficult. Even the AI field struggles here: an audit of dataset licensing found frequent errors on several major data hosting sites.
Rights are a four-stage data pipeline
Treat rights as a pipeline, and fix each stage.
Identify. Every asset gets a persistent, unique identifier that survives across systems, formats, and versions.
Link. Connect each asset to its creators, contributors, owners, and governing contracts. Rights are relationships, and relationships need to be modeled.
Express. Translate contract terms into structured, machine-readable fields: permitted uses, territories, windows, exclusivity, and whether AI training is allowed or reserved.
Prove. Keep an audit trail of provenance, changes, and approvals, so you can show who granted what, when, to anyone who asks.
Where it usually breaks
- Rights live in documents, not data. Nobody can query a contract folder.
- Catalogs are siloed by division, region, or acquisition, each with its own identifiers.
- Legacy asset systems weren't built to hold rights logic or talk to modern licensing and distribution tools.
- No single owner. Legal, content, and technology each assume another team is responsible.
- Metadata gets lost in distribution, stripped or overwritten on the way to platforms.
Almost all of these are data and systems issues. That makes them fixable.
What a working rights data layer looks like
- One unified metadata and rights repository on a modern cloud foundation, with consistent definitions across catalogs.
- AI-assisted enrichment that extracts terms from contracts and fills gaps, with human review on anything that carries legal weight.
- APIs into licensing, distribution, and ad systems, so rights checks happen automatically before content moves.
- An audit and provenance trail that can answer a regulator, partner, or AI licensee on demand.
Start with one decision
Don't try to clean the entire archive. Pick one high-value question, such as "Can we license this library for AI training?" Build the pipeline for that catalog, prove it, and extend it.
How Saguna helps
Saguna helps media and communications companies build the data foundation behind confident rights decisions. We unify fragmented catalogs on cloud platforms, apply AI and machine learning to extract and enrich rights data, and integrate it into the content, licensing, and distribution systems your teams already use. Where the operating model needs to change, our strategy and operations work gives rights data a clear owner and process.
The companies that monetize their content best in the AI era won't just own the best catalogs. They'll know exactly what they own.