DIGITIZATION AND CONTENT ENRICHMENT
Turning primary source archives into structured, discoverable content, at scale
Amnet pairs purpose-built AI with human review to move books, manuscripts, newspapers, and archival papers from scanned image to publication-ready XML, accurately and on schedule.
WHO WE ARE
Two decades of digitization, run like a controlled process
Amnet delivers OCR, Handwritten Text Recognition, XML creation, metadata enrichment, and content indexing across books, manuscripts, newspapers, periodicals, journals, and archival papers. An AI-assisted, human-verified workflow holds accuracy at scale without trading off quality.
25+
Years of digitization and content expertise
20-30
Active projects in the current pipeline
8
High-volume initiatives, each 100Ks of images
CORE SERVICES DELIVERED
Image Processing and QA
99.5% accuracy benchmark, verified at every batch.
Metadata Extraction and TEI XML
Structured, standards-compliant XML output.
Redaction Services
Sensitive data identified and removed to spec.
Controlled and Keyword Indexing
Taxonomy-aligned, AI-proposed, expert-reviewed.
Content Checking (UAT) and Sign-off
Formal review gate before delivery.
AI CAPABILITY
AI-powered capability highlights
Purpose-built AI layers embedded across the digitization pipeline, paired with human review where judgment matters most.
Image Cleanup
AI-assisted dust removal, deskewing, and orientation correction ahead of conversion.
Conversion
Automated TIFF/PDF-to-JPEG conversion at custom DPI per project specification.
OCR and HTR
Machine and handwritten text recognition with bounding-box content extraction.
Content Extraction
Structured metadata capture using authority files, TOCs, and catalogue data.
Article Split-Up
Automated segment- and article-level tagging within the TEI XML structure.
Controlled Indexing
AI-proposed taxonomy placement with confidence scoring, reviewed by experts.
SCOPE OF WORK
End-to-end processing of pre-digitized primary sources
Delivered as enriched, structured, quality-verified output, ready for platform ingestion.
Content Types Supported
Books and manuscripts
Periodicals: magazines, newspapers, comics and academic journals
Archive papers (e.g., government documents)
Input Formats
Digitised images: TIFF, JPEG, or PDF
Catalogue data: Excel or XML
Project-specific “spec” documents
Amnet Deliverables
Quality-checked, enhanced image files
TEI XML: OCR/HTR text, metadata, article-level tagging
Redacted images, indexed content and UAT/QA reports
METHODOLOGY
Proposed workflow and methodology
Seven controlled stages, each with a defined AI-assist point and a human quality checkpoint.
01
Input Analysis
Assess source quality; flag corrupt or incomplete files.
02
Image Conversion and QA
Convert to custom-spec JPEGs; cleanup dust, skew, orientation.
03
Metadata Extraction
Extract metadata from authority files and TOC.
04
OCR and Content Extraction
OCR/HTR extraction with content bounding boxes.
05
XML Conversion
OCR’ed content structured into standard XML + taxonomy.
06
XML Validation and QA
Automated validation and WYSIWYG preview QA.
07
Platform Ingestion and QA
Manifest ingested; content, image and metadata QA.
INDEXING TIERS
Content indexing tiers
Transparent depth options, indexed against your platform taxonomy, with AI carrying the load where confidence is highest.
Controlled Subject Indexing
Option A: 200–300 terms, up to 3 levels.
Option B: 400–500 terms, up to 5 levels.
AI-ASSISTED DELIVERY
AI proposes top-3 taxonomy placements with a confidence score from OCR’d text; broader Option A leans more on AI, deeper Option B routes more to human review.
Standard Metadata Indexing
Topic, content type, date/time period, geography, and person; person names resolved against LOC authority records.
AI-ASSISTED DELIVERY
Topic, type, date and location are auto-extracted with high confidence; ambiguous person-name matches route to human review.
Keyword Indexing
3–10 keywords assigned per item (issue or article) for fast, relevant discovery.
AI-ASSISTED DELIVERY
Most AI-assisted tier: an LLM ranks keyword candidates; a reviewer selects and finalizes the final list.
CASE STUDY
1 million pages, 16 weeks
A leading STEM and academic publisher: legacy scanned PDF book pages converted to client-specific XML.
The Challenge
- 1M legacy scanned image-PDF pages as source
- Right-file assessment across a 3.5K-file repository
- 16-week delivery window
- 2.6M references to be tagged per spec
- 100K math equations to MathML; 100K tables to XML
The Solution
- Centralized PDF-to-content extraction OCR engine
- Auto-transformation of XML per specification
- Auto-prediction of images and style application
- Human-in-the-loop WYSIWYG review tool
- Auto AI/ML reference tagging + automated Schematron validation
5→250
Titles/week scaled in 6 weeks
99.15%
Quality: only 17 of 2,004 titles re-evaluated
2K→24K
Pages/day capacity in 6 weeks
CASE STUDY
A 25-year partnership, sustained
Consistent OCR-to-XML delivery for a theological publisher, sustained across two decades of volume swings.
Extraction and Conversion Workflow
- Scan hardcopies to PDF
- OCR / keyboarding text extraction
- Metadata extraction
- XML / ePub conversion
- Proofreading and compare
OUTPUT: color images (300dpi PNG), Word text, plus XML / ePub
Fast-Track Team for Ad Hoc Requests. A dedicated expert team handles 5–10 short-turnaround titles every month, keeping ad hoc priorities on schedule.
WHY AMNET
A dependable partner for long-running archive programmes
AI + Human-Verified Accuracy
Quality assurance benchmarked at 99.5% accuracy, with structured exception reporting and rescan management.
Dedicated Project Management
A single PM as tactical lead, from kickoff to sign-off, SFTP setup, rolling delivery plans, and daily/weekly reporting.
Proven Scalability
Demonstrated ability to ramp capacity up or down, from 2,000 to 24,000 pages per day within weeks.
Long-Term Partnership
20+ year client relationships spanning STEM, academic, and theological publishing programmes.