DIGITIZATION AND CONTENT ENRICHMENT

Turning primary source archives into structured, discoverable content, at scale

Amnet pairs purpose-built AI with human review to move books, manuscripts, newspapers, and archival papers from scanned image to publication-ready XML, accurately and on schedule.

WHO WE ARE

Two decades of digitization, run like a controlled process

Amnet delivers OCR, Handwritten Text Recognition, XML creation, metadata enrichment, and content indexing across books, manuscripts, newspapers, periodicals, journals, and archival papers. An AI-assisted, human-verified workflow holds accuracy at scale without trading off quality.

25+

Years of digitization and content expertise

20-30

Active projects in the current pipeline​

8

High-volume initiatives, each 100Ks of images

CORE SERVICES DELIVERED

Image Processing and QA

99.5% accuracy benchmark, verified at every batch.

Metadata Extraction and TEI XML

Structured, standards-compliant XML output.

Redaction Services

Sensitive data identified and removed to spec.

Controlled and Keyword Indexing

Taxonomy-aligned, AI-proposed, expert-reviewed.

Content Checking (UAT) and Sign-off

Formal review gate before delivery.

AI CAPABILITY

AI-powered capability highlights

Purpose-built AI layers embedded across the digitization pipeline, paired with human review where judgment matters most.

Image Cleanup

AI-assisted dust removal, deskewing, and orientation correction ahead of conversion.

Conversion

Automated TIFF/PDF-to-JPEG conversion at custom DPI per project specification.

OCR and HTR

Machine and handwritten text recognition with bounding-box content extraction.

Content Extraction

Structured metadata capture using authority files, TOCs, and catalogue data.

Article Split-Up

Automated segment- and article-level tagging within the TEI XML structure.

Controlled Indexing

AI-proposed taxonomy placement with confidence scoring, reviewed by experts.

SCOPE OF WORK

End-to-end processing of pre-digitized primary sources

Delivered as enriched, structured, quality-verified output, ready for platform ingestion.

Content Types Supported

Books and manuscripts

Periodicals: magazines, newspapers, comics and academic journals

Archive papers (e.g., government documents)

Input Formats

Digitised images: TIFF, JPEG, or PDF

Catalogue data: Excel or XML

Project-specific “spec” documents

Amnet Deliverables

Quality-checked, enhanced image files

TEI XML: OCR/HTR text, metadata, article-level tagging

Redacted images, indexed content and UAT/QA reports

METHODOLOGY

Proposed workflow and methodology

Seven controlled stages, each with a defined AI-assist point and a human quality checkpoint.

01

Input Analysis

Assess source quality; flag corrupt or incomplete files.

02

Image Conversion and QA

Convert to custom-spec JPEGs; cleanup dust, skew, orientation.

03

Metadata Extraction

Extract metadata from authority files and TOC.

04

OCR and Content Extraction

OCR/HTR extraction with content bounding boxes.

05

XML Conversion

OCR’ed content structured into standard XML + taxonomy.

06

XML Validation and QA

Automated validation and WYSIWYG preview QA.

07

Platform Ingestion and QA

Manifest ingested; content, image and metadata QA.

INDEXING TIERS

Content indexing tiers

Transparent depth options, indexed against your platform taxonomy, with AI carrying the load where confidence is highest.

Controlled Subject Indexing

Option A: 200–300 terms, up to 3 levels.
Option B: 400–500 terms, up to 5 levels.

AI-ASSISTED DELIVERY

AI proposes top-3 taxonomy placements with a confidence score from OCR’d text; broader Option A leans more on AI, deeper Option B routes more to human review.

Standard Metadata Indexing

Topic, content type, date/time period, geography, and person; person names resolved against LOC authority records.

AI-ASSISTED DELIVERY

Topic, type, date and location are auto-extracted with high confidence; ambiguous person-name matches route to human review.

Keyword Indexing

3–10 keywords assigned per item (issue or article) for fast, relevant discovery.

AI-ASSISTED DELIVERY

Most AI-assisted tier: an LLM ranks keyword candidates; a reviewer selects and finalizes the final list.

CASE STUDY

1 million pages, 16 weeks

A leading STEM and academic publisher: legacy scanned PDF book pages converted to client-specific XML.

The Challenge

The Solution

5→250

Titles/week scaled in 6 weeks

99.15%

Quality: only 17 of 2,004 titles re-evaluated

2K→24K

Pages/day capacity in 6 weeks

CASE STUDY

A 25-year partnership, sustained

Consistent OCR-to-XML delivery for a theological publisher, sustained across two decades of volume swings.

Extraction and Conversion Workflow

OUTPUT: color images (300dpi PNG), Word text, plus XML / ePub

Fast-Track Team for Ad Hoc Requests. A dedicated expert team handles 5–10 short-turnaround titles every month, keeping ad hoc priorities on schedule.

WHY AMNET

A dependable partner for long-running archive programmes

AI + Human-Verified Accuracy

Quality assurance benchmarked at 99.5% accuracy, with structured exception reporting and rescan management.

Dedicated Project Management

A single PM as tactical lead, from kickoff to sign-off, SFTP setup, rolling delivery plans, and daily/weekly reporting.

Proven Scalability

Demonstrated ability to ramp capacity up or down, from 2,000 to 24,000 pages per day within weeks.

Long-Term Partnership

20+ year client relationships spanning STEM, academic, and theological publishing programmes.

Explore Our Digital Accessibility Expertise

Explore our editorial expertise

Our eLearning Expertise Designed for Modern Businesses ​

Search

Our organization is committed to making our website accessible to everyone, including people with disabilities. We are pleased to report that the majority of our website is fully accessible and meets the Web Content Accessibility Guidelines (WCAG) 2.1 Level AA. These guidelines provide a set of international standards for web accessibility and outline the requirements for making web content more accessible to people with disabilities.

However, we recognize that the chat bot is currently not fully accessible and we are working to improve its accessibility. We apologize for any inconvenience this may cause and are committed to replace the chat bot with an accessible one.

If you have any questions or need assistance accessing any content on our website, please contact us at [email protected]. We will do our best to provide the information to you in an alternative format.
Thank you for your understanding.