
Most companies don’t have a data shortage.
They have a context shortage.
Industry research puts it starkly: 68% of available business data goes unused, and nearly half of employees say they struggle to find the documents they need. That gap doesn’t go away when AI enters the picture. It gets worse, because AI can’t tell the difference between good data and undocumented data. It just uses what it’s given.
DataLab Group felt that gap firsthand. It’s Crédit Agricole’s center of expertise for designing industrial-grade, innovative AI solutions across the entire group, 145,000 people and 53 million customers strong.
Since 2022, they’ve been on a certification and labeling path to prove that AI solutions can be innovative, trustworthy, and responsible at the same time. And every model they ship eventually faces that certification board.
The question was never whether the AI would work. It was whether anyone could explain why. You can’t certify what you can’t explain. And you can’t explain data nobody documented.
DataGalaxy is the platform that closed that gap.
Challenges
DataLab Group’s AI runs on three very different kinds of data:
- Analytical AI reads structured banking data
- Document AI reads scanned contracts and forms
- Textual AI reads free text, like emails and reports
That mix created real friction:
- Documentation was fragmented across teams. No shared way to describe a dataset once and reuse it everywhere.
- Structured, semi-structured, and unstructured data each needed different handling, but teams needed one system, not three.
- Every indicator had to trace back to its source, sometimes across multiple systems, for certification to hold up.
- Consistency got harder as more projects and teams got involved.
- None of this could slow a team that ships AI for a living.
The data existed. It just had no context: what it meant, where it came from, how it could be used. Without that, AI couldn’t scale responsibly, no matter how good the models were.
Why DataGalaxy
DataLab Group picked DataGalaxy on three criteria:
- It had to handle every data type they worked with, structured, semi-structured, and unstructured alike
- It had to adapt to each project, not force one mold onto every use case
- It had to be simple enough that teams would actually use it, day in and day out
A responsive support team sealed the decision.
“"I'd recommend DataGalaxy for three main reasons: ease of use, adaptability to different use cases, and efficient, available support teams."”
That flexibility let one platform document analytical, document, and textual AI consistently, without forcing three different problems into one format.
What context looks like in practice
For DataLab Group, giving data context meant more than writing a description in a text box. It meant capturing, for every asset:
- What it is — a clear definition, in language both data teams and business teams understand
- Where it came from — the source system, and every transformation applied along the way
- Who owns it — a named steward accountable for keeping the documentation current
- How it’s allowed to be used — usage licenses and constraints, especially for open data shared across teams
- How it connects to an AI indicator — the link between a raw dataset and the model output built on top of it
That’s the difference between a data catalog that just lists what exists and one that explains what it means.
How it came together
DataLab Group didn’t try to document everything at once. They piloted first.
Step 1: Pilot three use cases. The team retro-documented three AI use cases as part of their certification process, one from each major family, analytical, document, and textual AI.
Step 2: Turn the pilot into templates. That pilot became three reusable documentation templates, one each for structured data, text, and images, all sharing a common foundation. Every future project could start from the same place instead of reinventing its own approach.
Step 3: Generalize the practice. DataGalaxy became the system for structuring raw data, indicators, and derived assets consistently, and for capturing ownership, definitions, and usage before any dataset entered an AI pipeline.
Step 4: Extend it beyond AI. The same foundation now supports an Open Data Mart, where open data is retrieved, reworked, and made available to every project team instead of being rebuilt from scratch each time someone needs it.
Today, DataGalaxy runs across DataLab Group’s AI projects, including those going through formal certification and labeling. The team now tracks around 15 different data sources in the platform, along with their documentation and usage licenses. And they can trace, in detail, how 250 raw data points get transformed into 400 indicators, sometimes pulling from multiple sources at once. That’s the kind of trail a certification board actually wants to see.
DataLab Group’s own advice to teams starting out: don’t try to boil the ocean. Start with a minimum viable version to cover the first needs, then reinforce it step by step, building on that foundation. And budget real time for it. Setting up the tool properly, and keeping it fed with good documentation, takes more effort than people expect.
Outcomes
- 100+ AI-ready data assets and indicators documented and structured
- ~15 data sources centralized, documented, and licensed for reuse
- 250 raw data points transformed into 400 indicators, fully traceable, sometimes across multiple sources at once
- Multiple AI data types covered by one framework: structured data, text, documents, and images
- 3 use cases retro-documented in the pilot, producing 3 reusable templates now used group-wide
- End-to-end traceability, from raw data to AI indicators, backing certification for every labeled project
- An Open Data Mart, distributing open data across the organization instead of every team sourcing it alone
- 1 single source of truth for current and future data and AI initiatives
“"We're rolling the tool out across every project at DataLab Group. We run a lot of AI projects that follow our labeled and certified methodology, and project documentation in DataGalaxy is now part of that process."”
The takeaway
Trusted AI runs on context, not just data.
DataLab Group didn’t need more models. It needed every model to come with a clear, structured account of the data behind it, one an auditor, a business team, or the next AI project could actually use.
That’s what turned documentation from a bottleneck into the backbone of how DataLab Group scales AI it can stand behind. And it’s the same principle behind every certification the group has earned since: govern the data, and the AI built on top of it earns trust by default.




