AI Data Preparation Platform

AI-powered data preparation for trusted and sovereign AI.

Datahunter turns unstructured documents into metadata-rich, searchable, and analysis-ready data, deployed on your infrastructure, aligned to your standards.

92%Time savings reported by NRC researchers
Protected B-readyOn-premises, no cross-border data flows
Real GC dataBuilt on real department glossaries and official translations, no AI slop

What is Datahunter?

Datahunter is an AI-powered platform that helps organizations prepare high-quality data for trusted and sovereign AI initiatives. Many valuable datasets exist only within unstructured documents — scientific publications, reports, and technical materials. Datahunter automatically identifies relevant sources, detects PII, applies metadata and security classification, extracts structured data, and organizes it into metadata-rich, searchable, and analysis-ready formats. This significantly reduces the time required to build datasets for analytics, RAG knowledge systems, and AI inference.

What Datahunter does

Four steps from documents to datasets

Each step runs on your infrastructure, with full traceability from source document to structured output.

FindFind

Identifies relevant sources across scientific repositories, document stores, and technical materials — surfacing what exists across formats and siloes.

ClassifyClassify

Detects PII, applies metadata and security classification automatically according to your organizational context and standards.

ExtractExtract

Pulls structured data out of unstructured documents — scientific papers, reports, technical materials — and refines results with AI-assisted query tools.

OrganizeOrganize

Builds searchable, analysis-ready outputs for analytics pipelines, RAG knowledge systems, and AI inference — integrated with your existing infrastructure and workflows.

National Research Council Canada

National Research Council Canada

Federal Case Study

“
It would take a researcher 30 to 60 minutes to read and extract LCI data from a publication. When using Datahunter, the researcher would spend about 5 minutes per publication.

— Cyrille D., National Research Council Canada

NRC reports up to 92% time-savings for researchers using Datahunter

Extraction and classification logic developed and validated against real research data, refined in partnership with Canada's federal science and research community.

Sovereign by design

On-premises, Protected B-capable

Self-contained deployment with no mandatory SaaS dependency. Configurable to department-specific workflows rather than forcing adaptation to the tool.

  • Fully on-premises deployment — Protected B-capable
  • No multi-tenant SaaS dependency
  • Configurable to department-specific workflows
  • No cross-border data flows
  • No mandatory external service dependencies
Built for compatibility

Integrates with your existing stack

Combines local models with deterministic validation and open-source document parsing. Every extraction step produces results you can trace and verify — not a black box.

  • Out-of-the-box DSpace integration
  • Termium terminology alignment (bilingual standards)
  • Integrates with existing infrastructure and workflows
  • Local models + deterministic validation — fully auditable
  • Open-source document parsing, no black-box outputs
DSpace

Integrates with DSpace

Out-of-the-box compatibility with DSpace so Datahunter works with your existing research repository infrastructure — no migration required.

Learn how Datahunter can accelerate your data maturity.

Request a Demo

Commercialization

Available through Pathways to Commercialization (PTC)

Datahunter is available through the NRC's Pathways to Commercialization program, connecting federal science organizations with innovative Canadian technology to accelerate their AI and data capabilities.

Get in touch
Loading analytics...