Trustpilot’s AI-ready data catalog starts with ownership—not chat
A five-step rollout across 500,000-plus assets shows why reliable data agents need governed context before they need a conversational interface.
Trustpilot’s route toward enterprise data agents began with a less glamorous task: establishing who owns the data.
A DataHub customer account describes an estate of more than 500,000 assets accumulated over 19 years across AWS and Google Cloud. Before the catalog rollout, finding data and tracing dependencies meant searching GitHub and Slack or asking people who held the history in their heads. Trustpilot’s answer was a five-pillar governance sequence: ownership, classification, metadata, lineage and data quality.
Sequence the context before the agent
The order is the useful part for teams planning natural-language analytics. Trustpilot made ownership the first requirement because descriptions, classifications and operational context need accountable maintainers. It then added criticality classification, initially narrow table- and column-level descriptions, cross-platform lineage and quality assertions.
That sequence turns “AI-ready data” into an implementation checklist rather than a model claim. An agent can retrieve a table name without knowing whether the asset is authoritative, sensitive, maintained or downstream of a planned migration. Trustpilot’s catalog records are intended to supply those missing decisions.
The rollout also separates discovery from endorsement. The company has nearly 500 Looker dashboards, according to the case study, and marks selected dashboards as governed sources of truth. DataHub’s browser extension exposes documentation, ownership and trust status inside Looker, so users do not have to interpret every discovered dashboard as equally reliable.
Lineage is an agent input, not just a diagram
The strongest implementation detail is how Trustpilot uses lineage during a migration from self-hosted MongoDB on AWS to Amazon DocumentDB. Engineers can inspect downstream impact in DataHub’s interface, query it through chat, or pull lineage context into Claude Code and Copilot through DataHub’s MCP integration.
That makes lineage operational context for an assistant: not merely “where is the table?” but “what will this change affect, and who should review it?” For teams exposing catalog context through MCP, the case suggests a practical boundary: connect the agent only after ownership and dependency records are useful enough to support a human decision.
What the case study does not prove
This is a vendor-published customer account, not an independent evaluation. It does not report text-to-SQL accuracy, agent task-completion rates or a measured reduction in hallucinations. It also says PII classification is still a planned next step while separately describing governance and compliance fundamentals as anchored in the catalog.
So the defensible takeaway is narrower than “a catalog makes agents reliable.” Trustpilot has built the context layer it expects future agents to use, and it did so by treating ownership, trusted assets and lineage as prerequisites. Teams following the pattern should measure the agent separately—but they should not expect a chat layer to repair missing governance underneath it.
sources
comments · 0