AWS Database Blog

Amagi’s intelligent media operations with Amazon Neptune

This post was co-written with Ajith N N, Sr Architect at Amagi Media Labs.

When Amagi set out to build a unified metadata store for its intelligent media operations across over 2,500 channels and over 150 countries, the challenge was clear: traversing billions of interconnected relationships without performance degradation. Relational databases couldn’t scale to this depth, so Amagi chose a graph database.

In the modern media landscape, the volume of content and the complexity of its associated metadata continue to grow. Media companies no longer only manage video files. They manage a web of interconnected data points including multilingual credits, regional licensing rights, and complex version lineages. Relational databases are not optimized for highly connected datasets. Linking a series to its seasons, episodes, and trailers creates deep relationship chains. In SQL databases, each additional layer adds join operations that degrade query performance.

To address this challenge, Amagi built the Global Metadata Store (GMS), a unified intelligence layer for the media supply chain. The GMS uses a Resource Description Framework (RDF) knowledge graph on Amazon Neptune to model billions of relationships. The graph supports deep semantic queries while maintaining sub-500ms response times for operational workloads.

In this post, we describe the key technical challenge, blank node synchronization, that Amagi encountered when synchronizing graph data to Amazon OpenSearch Service, and the hybrid query strategy that balances depth with speed.

Challenges with previous architecture

Amagi is a cloud-based media technology company serving over 800 content brands across the globe. The company’s platform powers the entire media supply chain, from content preparation and playout to distribution and monetization. Media metadata is hierarchical and relational by nature. Consider a television series: a single brand connects to multiple series, each containing seasons, and each season containing episodes. Every episode may have associated trailers, teasers, and promotional clips. Live events generate recordings and AI-generated highlights that must link back to the parent event. Distribution rights and localized audio tracks attach to specific versions across geographic regions.

Relational database: Query complexity

Entity-relationship diagram of seven media tables joined by foreign keys to resolve a Season 3 licensing query

Figure 1: In a relational database, resolving media relationships requires multiple expensive joins that degrade performance as relationship depth grows

When this data lives in relational tables, discovering relationships requires traversing multiple layers through joins. Consider a query such as “find all promotional assets for Season 3 episodes licensed in Germany.” In a relational database, this might span five or six tables. As the dataset grows to billions of records, these join operations become increasingly expensive.

The same media hierarchy modeled as an RDF graph in Amazon Neptune, with nodes traversed directly along labeled edges

Figure 2: The same relationships modeled as an RDF graph in Amazon Neptune, where queries traverse edges directly instead of computing joins

Amagi needed a solution that could scale to billions of metadata records while supporting deep, multi-hop semantic queries without performance degradation. The system also needed to serve high-frequency transactional reads with consistent sub-second latency for scheduling and distribution APIs.

Solution overview

The GMS uses a polyglot persistence architecture that pairs two managed services. Amazon Neptune, a fully managed graph database service, handles deep graph traversal. Amazon OpenSearch Service, a managed service for deploying and scaling OpenSearch clusters in the AWS Cloud, handles high-throughput transactional reads.

AWS services used

  • Neptune serves as the source of truth for the knowledge graph, handling deep search queries, lineage tracking, and complex rights inheritance across the graph using RDF and SPARQL.
  • High-frequency transactional reads with sub-500ms latency come from OpenSearch Service, which serves millions of requests for scheduling and distribution APIs.
  • Amazon Elastic Kubernetes Service (Amazon EKS) hosts the custom synchronization application that bridges Neptune and OpenSearch.
Amagi Global Metadata Store architecture: Amazon API Gateway, Amazon EKS pods, Amazon Neptune, and Amazon OpenSearch Service inside a virtual private cloud

Figure 3: The GMS architecture, where Amazon EKS hosts the sync application, Amazon Neptune serves as the source of truth, and Amazon OpenSearch Service provides the fast read layer

The blank node synchronization challenge

Synchronizing blank nodes from Neptune to OpenSearch Service is the core technical problem the GMS solves. The following sections explain what blank nodes are, why they resist standard synchronization, and how Amagi’s custom sync layer resolves them.

Understanding blank nodes

In an RDF graph, most nodes have a Uniform Resource Identifier (URI) that makes them globally unique and addressable. However, media metadata often requires groupings of information that don’t need a formal identity, such as coordinates for an image crop, a complex rights restriction, or a nested set of talent credits. Blank nodes (or anonymous nodes) serve this purpose. They act as join points for multiple properties without requiring a formal URI.

For example, instead of stating that an actor appears in a movie, a blank node can connect the actor to a specific role, character name, and credit order, all bundled as a single entity within the graph.

Resolving blank node sync for nested media metadata

AWS provides native integration to sync Neptune data to OpenSearch through Neptune Streams and an AWS Lambda-based reloader. However, this approach faces a significant challenge with blank nodes:

  • Non-persistent identifiers: Blank node IDs in Neptune aren’t persistent across exports or certain internal operations.
  • Context fragmentation: Neptune Streams capture changes at the statement (triple) level. A stream event for a blank node often lacks the global context of the parent URI it belongs to.
  • Sync incompatibility: The native OpenSearch sync expects stable IDs for upserts. The transient nature of blank node IDs makes it impossible to reliably map nested metadata to the correct search document using standard tools.

The custom sync solution

To bridge this gap, Amagi built a custom synchronization application. The application intercepts Neptune Stream events and adds a resolution layer. This layer traverses the graph to map blank nodes back to their nearest named parent URI. The application reassembles anonymous triples into structured JSON fragments, flattens the data, and indexes it into OpenSearch Service.

This approach helps keep complex nested metadata (originally stored as anonymous nodes) searchable with sub-500ms latency while preserving the rich context of the knowledge graph.

The hybrid query strategy of GMS

The GMS routes queries to the appropriate backend based on their characteristics:

  • Neptune handles deep search queries requiring multi-hop traversal, lineage tracking, and complex rights inheritance. These queries benefit from the graph’s native ability to traverse relationships without joins.
  • OpenSearch Service handles high-frequency transactional reads where speed is critical. By indexing resolved blank nodes, the GMS serves millions of concurrent read requests, loading complex series pages with sub-500ms latency.

Worked example: Routing one query across Neptune and OpenSearch

This uses a subset of the real GMS graph in Amagi’s EBUCore-based model (ns1 = EBUCore, ns3 = the Amagi extension). The request: for a series, find episodes licensed for distribution in India (IN) and return their ad-break cue points. Note that CoverageRestrictions, Cuepoint and Identifier are all anonymous (blank) nodes, precisely the nodes the sync layer must resolve to a named parent.

1. Source graph in Neptune (Turtle subset). The rights grant, cue points, and identifier have no URI of their own:

@prefix ns1: <http://www.ebu.ch/metadata/ontologies/ebucore/ebucore#> .
@prefix ns3: <https://metadata.amagi.tv/ebu/amagi_ebucoreExtension#> .
@prefix iso: <http://www.ebu.ch/metadata/ontologies/skos/ebu_Iso3166_CountryCodeCS#> .
@prefix ent: <https://compass.amagi.tv/gms/entity/> .

ent:asset_id_2 a ns1:Episode ;
ns1:episodeNumber "2" ; ns1:title "Feeling Loving"@en ;
ns1:isEpisodeOfSeries ent:series_id_1 ;
ns1:isEpisodeOfSeason ent:season_id_1 ;
ns1:hasIdentifier [ a ns1:Identifier ; # blank node
ns1:hasIdentifierType "urn:amagi:publisher:assetId" ;
ns1:identifierValue "da9cd6d49928e78ad08c1706220" ] ;
ns1:hasCoverageRestrictions [ a ns1:CoverageRestrictions ; # blank node
ns1:rightsTerritoryIncludes iso:_CA, iso:_IN, iso:_US ;
ns1:rightsTerritoryExcludes iso:_AS, iso:_MX ;
ns1:rightsStartDateTime "2015-01-01T00:00:00+00:11" ;
ns1:rightsEndDateTime "2035-01-01T00:00:00+00:11" ] ;
ns3:hasCuepoint [ a ns3:Cuepoint ; ns3:cuepointTimestamp "1084.734" ; # blank
ns3:publisherCuepointOffset "00:18:04;23" ; ns3:hasCuepointType "ad_break" ] ,
[ a ns3:Cuepoint ; ns3:cuepointTimestamp "1731.467" ;
ns3:publisherCuepointOffset "00:28:51;14" ; ns3:hasCuepointType "ad_break" ] ,
[ a ns3:Cuepoint ; ns3:cuepointTimestamp "2253.234" ;
ns3:publisherCuepointOffset "00:37:33;07" ; ns3:hasCuepointType "ad_break" ] .

2. Deep path → Neptune (SPARQL). Multi-hop traversal into two different blank-node structures (coverage and cue points) — the shape SQL joins handle poorly and the graph handles natively:

PREFIX ns1: <http://www.ebu.ch/metadata/ontologies/ebucore/ebucore#>
PREFIX ns3: <https://metadata.amagi.tv/ebu/amagi_ebucoreExtension#>
PREFIX iso: <http://www.ebu.ch/metadata/ontologies/skos/ebu_Iso3166_CountryCodeCS#>
PREFIX ent: <https://compass.amagi.tv/gms/entity/>

SELECT ?episode ?cueOffset WHERE {
?episode ns1:isEpisodeOfSeries ent:series_id_1 ;
ns1:hasCoverageRestrictions ?cov ;
ns3:hasCuepoint ?cue .
?cov ns1:rightsTerritoryIncludes iso:_IN . # licensed in India
?cue ns3:hasCuepointType "ad_break" ;
ns3:cuepointTimestamp ?cueOffset .
} ORDER BY ?cueOffset

3. The custom sync application resolves those blank nodes to their parent asset URI and flattens the subgraph into a single OpenSearch document:

{
"asset_id": "https://compass.amagi.tv/gms/entity/asset_id_2",
"type": "Episode", "episode_number": 2, "title": { "en": "Feeling Loving" },
"series_id": "https://compass.amagi.tv/gms/entity/series_id_1",
"season_id": "https://compass.amagi.tv/gms/entity/season_id_1",
"publisher_asset_id": "da9cd6d49928e78ad08c1706220",
"coverage": {
"territory_includes": ["CA", "IN", "US"],
"territory_excludes": ["AS", "MX"],
"rights_start": "2015-01-01T00:00:00+00:11",
"rights_end": "2035-01-01T00:00:00+00:11"
},
"ad_breaks": [
{ "offset": 1084.734, "tc": "00:18:04;23" },
{ "offset": 1731.467, "tc": "00:28:51;14" },
{ "offset": 2253.234, "tc": "00:37:33;07" }
]
}

4. Fast path → OpenSearch. The high-frequency asset/series page load becomes a single-document term fetch keyed on the asset URI — no traversal:

GET gms-assets/_search
{ "query": { "term": { "asset_id":
"https://compass.amagi.tv/gms/entity/asset_id_2" } } }

Routing rule: ad-hoc, lineage, and rights-inheritance queries (unknown shape, arbitrary depth) go to Neptune. The flattened per-asset document above serves known-entity page loads and high-concurrency filters from OpenSearch. Writes land in Neptune first, then the sync layer resolves the blank CoverageRestrictions, Cuepoint, and Identifier nodes to the asset URI and upserts the document.

Trade-off: A denormalized read cache and eventual consistency

The flattened OpenSearch document is a denormalized cache of a subgraph, and that is the design tension worth naming. When rights change at the series level, Neptune recomputes inheritance instantly. However, the sync layer must update every affected flattened document before the change is visible on the fast path. The “sub-500ms” figure is therefore a read-latency service-level agreement, not the write-to-visible propagation time.

In practice this creates two consistency domains. Neptune is strongly consistent and authoritative for correctness-critical checks (such as whether an asset can air in a given territory). OpenSearch is eventually consistent and optimized for read volume. Correctness-sensitive reads should either query Neptune directly or tolerate the sync lag explicitly. Treating the search index as authoritative would risk serving a stale rights state during the propagation window.

Results and benefits

Amagi’s hybrid design combines the structural intelligence of a graph with the sub-500ms retrieval of a search index. It delivers several key capabilities.

The GMS delivers measurable performance at scale. In production, the system manages over 500 million triples representing more than one million media assets. It sustains 300–500 read requests per second, with peaks exceeding 1,000 requests per second. Latency stays under 500ms for transactional reads. Compared to the previous approach, query performance improved by approximately 30 percent. When distribution rights change at the series level, every localized clip, trailer, and highlight automatically inherits the correct restrictions through graph lineage. This removes the need to manually reconcile hundreds of derived assets.

The architecture also delivers two lasting operational advantages. First, with the RDF-based model, Amagi can ingest metadata from studios and AI services without database migrations. New types, such as AI-generated scene sentiment or viewer engagement signals, integrate without schema changes. Second, for live sports and news workflows, the custom sync application orchestrates simultaneous writes. As a live event records, highlights connect to the parent event in Neptune and become immediately searchable in OpenSearch for global distribution.

“Our strategic partnership with AWS has been instrumental in reimagining how media assets are managed at a global scale. By migrating to Neptune, we’ve moved beyond the performance limitations of traditional relational joins, allowing us to traverse complex hierarchies of brands, series, and rights in milliseconds. This shift not only accelerates our content supply chain but also helps our customers to monetize their libraries with exceptional speed and agility.”

— Srividya Srinivasan, CTO, Amagi Media Labs

Cross-domain applicability

The blank node synchronization pattern that Amagi developed extends beyond media to any industry modeling complex, nested relationships in RDF-based knowledge graphs. This architecture provides a blueprint for combining the semantic richness of RDF with the performance requirements of modern applications across healthcare, finance, and supply chain domains.

What’s next

Amagi continues to expand the GMS capabilities. Planned enhancements include machine learning (ML)-powered metadata enrichment, which is more straightforward because of the extensibility of Amazon Neptune, and sharded clusters that scale out while still supporting deep queries with federation.

Conclusion

Amagi’s Global Metadata Store shows how a Neptune-powered knowledge graph, paired with a custom synchronization layer and OpenSearch Service, can address metadata challenges in modern media operations. By keeping metadata connected, enriched, and actionable, Amagi provides a foundation for intelligent media supply chains.

To get started with Neptune for your own graph workload or to learn more about knowledge graphs, see the Amazon Neptune documentation.


About the authors

Ajith N N

Ajith N N

Ajith is a Director of Engineering at Amagi. He has over 12 years of experience in data management systems across both media and retail industries.

Smiti Guru

Smiti Guru

Smiti is an AWS Senior Solutions Architect specializing in database, analytics, AI, and generative AI solutions. She works with customers to architect scalable, intelligent systems that transform business operations through cloud technologies.

Sanjay Singh

Sanjay Singh

Sanjay is a Senior Solutions Architect at AWS India with nearly 22 years of expertise in Media & Entertainment, specializing in enterprise architecture, platform engineering, and cloud-driven digital transformation.