<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>OCR vs HTR on File Format Blog</title>
    <link>https://blog.fileformat.com/tag/ocr-vs-htr/</link>
    <description>Recent content in OCR vs HTR on File Format Blog</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Wed, 07 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.fileformat.com/tag/ocr-vs-htr/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>How to Choose the Right File Format for Handwritten Text Recognition (HTR)</title>
      <link>https://blog.fileformat.com/ocr/ocr-file-formats-for-historical-and-handwritten-documents/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      
      <guid>https://blog.fileformat.com/ocr/ocr-file-formats-for-historical-and-handwritten-documents/</guid>
      <description>Discover the best OCR and HTR file formats for historical manuscripts and handwritten archives, comparing ALTO, PAGE XML, hOCR, and TEI-XML.</description>
      <content:encoded><![CDATA[<p><strong>Last Updated</strong>: 07 Oct, 2026</p>
<figure class="align-center ">
    <img loading="lazy" src="images/ocr-file-formats-for-historical-and-handwritten-documents.png#center"
         alt="OCR File Formats for Historical and Handwritten Documents"/> 
</figure>

<h2 id="ocr-file-formats-for-historical-and-handwritten-documents">OCR File Formats for Historical and Handwritten Documents</h2>
<p>Preserving cultural heritage through digitization has entered a renaissance. While early optical character recognition (OCR) systems were engineered to parse pristine, machine-printed twentieth-century paperwork, modern cultural heritage institutions face a far messier, richer challenge: medieval manuscripts, nineteenth-century cursive correspondence, brittle illuminated folios, and century-old registries.</p>
<p>Transcribing historical documents is no longer just about extracting plain ASCII text. It requires capturing <strong>context</strong>—the physical geometry of the page, baseline curves, bleed-through corrections, marginalia, abbreviations, and paleographic uncertainty. This specialized field, frequently categorized under Handwritten Text Recognition (HTR), depends heavily on the file format used to store and exchange both image coordinates and textual layers.</p>
<p>Selecting the wrong schema can strip away vital baseline data, break alignment with IIIF (International Image Interoperability Framework) manifests, or prevent long-term digital preservation. Here is a definitive guide to the leading OCR and HTR file formats for historical documents, their structural strengths, and how to determine the right choice for your archival pipeline.</p>
<h2 id="the-historical-dilemma-why-plain-text15-and-standard-pdfs1-fail">The Historical Dilemma: Why <a href="https//docs.fileformat.com/word-processing/txt/">Plain Text</a> and Standard <a href="https://docs.fileformat.com/pdf/">PDFs</a> Fail</h2>
<p>Machine-printed OCR often outputs simple <code>.txt</code> files or &ldquo;sandwich&rdquo; PDFs with hidden text layers beneath the scan. For historical manuscripts and cursive handwriting, these outputs fail for three core reasons:</p>
<ol>
<li><strong>Non-Linear Text and Complex Layouts:</strong> Historical scribes did not adhere to neat rectangular grids. Text cascades into margins, wraps around illuminated initials, weaves between inserted interlinear corrections, or runs vertically along the spine.</li>
<li><strong>Curved and Slanted Baselines:</strong> Cursive writing rarely follows a rigid horizontal axis. HTR engines like Transkribus, Kraken, and eScriptorium rely on polyline baselines rather than bounding boxes to interpret ligature-heavy scripts.</li>
<li><strong>Paleographic Complexity and Metadata:</strong> Archival research requires tracking abbreviations, historical spelling variations, damaged readings, and line-level confidence scores. Standard document formats discard this granularity.</li>
</ol>
<p>To maintain fidelity to the original artifact, the archival community relies on structured XML schemas designed to preserve layout topology alongside transcribed text.</p>
<h2 id="1-page-xml-the-gold-standard-for-handwritten-text-recognition-htr">1. PAGE XML: The Gold Standard for Handwritten Text Recognition (HTR)</h2>
<p>Developed by the PRImA (Pattern Recognition &amp; Image Analysis) Research Lab, <strong>PAGE XML</strong> (Page Analysis and Groundtruth Elements) is widely considered the state-of-the-art format for handwritten text recognition and advanced layout analysis.</p>
<h3 id="core-architecture">Core Architecture</h3>
<p>PAGE XML treats the physical document as a hierarchical structure:</p>
<ul>
<li><code>PcGts</code> (Root)
<ul>
<li><code>Page</code> (Image dimensions and overall reading order)
<ul>
<li><code>TextRegion</code> (Paragraphs, headings, marginal notes, catchwords)
<ul>
<li><code>TextLine</code>
<ul>
<li><code>Coords</code> (Polygon coordinates around the line)</li>
<li><code>Baseline</code> (A series of points following the true baseline of the script)</li>
<li><code>TextEquiv</code> (The recognized text, with optional confidence metrics)</li>
</ul>
</li>
</ul>
</li>
</ul>
</li>
</ul>
</li>
</ul>
<h3 id="why-it-excels-for-historical-manuscripts">Why It Excels for Historical Manuscripts</h3>
<ul>
<li><strong>Polygon and Polyline Precision:</strong> Instead of forcing characters into rectangular bounding boxes, PAGE XML uses multi-point polygon boundaries and continuous baselines. This prevents overlapping cursive ascenders and descenders from interfering with segmentation.</li>
<li><strong>Granular Structural Types:</strong> Regions can be classified precisely (e.g., <code>marginalia</code>, <code>drop-capital</code>, <code>signature-mark</code>, <code>header</code>, <code>editorial-note</code>).</li>
<li><strong>Broad Software Ecosystem:</strong> It serves as the primary internal and export schema for flagship HTR platforms like <strong>Transkribus</strong>, <strong>eScriptorium</strong>, and <strong>Kraken</strong>.</li>
</ul>
<h2 id="2-alto-xml-the-library-and-archival-powerhouse">2. ALTO XML: The Library and Archival Powerhouse</h2>
<p><strong>ALTO (Analyzed Layout and Text Object)</strong> is an open XML standard maintained by the Library of Congress and widely adopted by national libraries, including the Bibliothèque nationale de France (BnF) and the British Library.</p>
<h3 id="core-architecture-1">Core Architecture</h3>
<p>ALTO structures layout hierarchically from <code>Page</code> to <code>PrintSpace</code>, down to <code>TextBlock</code>, <code>TextLine</code>, and <code>String</code> (individual words or tokens). It is frequently packaged inside a <strong>METS</strong> (Metadata Encoding and Transmission Standard) wrapper to associate structural metadata with high-resolution master images.</p>
<h3 id="key-strengths-and-use-cases">Key Strengths and Use Cases</h3>
<ul>
<li><strong>Mass Digitization Workflows:</strong> ALTO was designed with industrial-scale newspaper and book digitization in mind. It cleanly encodes font attributes, word-level coordinates, character confidence, and white spaces.</li>
<li><strong>Modern HTR Support (ALTO 4):</strong> Earlier versions of ALTO leaned heavily on rectangular coordinates (<code>HPOS</code>, <code>VPOS</code>, <code>WIDTH</code>, <code>HEIGHT</code>). However, starting with <strong>ALTO version 4</strong>, the schema introduced <code>&lt;Shape&gt;</code> polygons and polyline baselines, closing the functional gap with PAGE XML for handwritten materials.</li>
<li><strong>Long-Term Preservation:</strong> Because it is an official standard backed by international library consortia, ALTO guarantees long-term stability and archive-grade backwards compatibility.</li>
</ul>
<h2 id="3-hocr-the-web-first-lightweight-standard">3. hOCR: The Web-First, Lightweight Standard</h2>
<p>Created by Thomas Breuel, <strong>hOCR</strong> takes a pragmatic approach: rather than creating an entirely new XML schema, it embeds layout and transcription metadata directly into semantic HTML/XHTML using microformats and class attributes.</p>
<h3 id="typical-syntax-example">Typical Syntax Example</h3>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-html" data-lang="html"><span style="display:flex;"><span>&lt;<span style="color:#f92672">div</span> <span style="color:#a6e22e">class</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;ocr_page&#34;</span> <span style="color:#a6e22e">id</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;page_1&#34;</span> <span style="color:#a6e22e">title</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;bbox 0 0 2480 3508&#34;</span>&gt;
</span></span><span style="display:flex;"><span>  &lt;<span style="color:#f92672">div</span> <span style="color:#a6e22e">class</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;ocr_carea&#34;</span> <span style="color:#a6e22e">id</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;block_1_1&#34;</span>&gt;
</span></span><span style="display:flex;"><span>    &lt;<span style="color:#f92672">p</span> <span style="color:#a6e22e">class</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;ocr_par&#34;</span> <span style="color:#a6e22e">id</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;par_1_1&#34;</span>&gt;
</span></span><span style="display:flex;"><span>      &lt;<span style="color:#f92672">span</span> <span style="color:#a6e22e">class</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;ocr_line&#34;</span> <span style="color:#a6e22e">id</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;line_1_1&#34;</span> <span style="color:#a6e22e">title</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;bbox 150 320 2200 410; baseline 0 -5&#34;</span>&gt;
</span></span><span style="display:flex;"><span>        &lt;<span style="color:#f92672">span</span> <span style="color:#a6e22e">class</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;ocrx_word&#34;</span> <span style="color:#a6e22e">id</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;word_1_1&#34;</span> <span style="color:#a6e22e">title</span><span style="color:#f92672">=</span><span style="color:#e6db74">&#34;bbox 150 325 380 405; x_wconf 92&#34;</span>&gt;Incipit&lt;/<span style="color:#f92672">span</span>&gt;
</span></span><span style="display:flex;"><span>      &lt;/<span style="color:#f92672">span</span>&gt;
</span></span><span style="display:flex;"><span>    &lt;/<span style="color:#f92672">p</span>&gt;
</span></span><span style="display:flex;"><span>  &lt;/<span style="color:#f92672">div</span>&gt;
</span></span><span style="display:flex;"><span>&lt;/<span style="color:#f92672">div</span>&gt;
</span></span></code></pre></div><h3 id="pros-and-cons-for-historical-materials">Pros and Cons for Historical Materials</h3>
<ul>
<li><strong>Pros:</strong> Browsers can render it natively. It is easily transformed using vanilla CSS and JavaScript, and it is the default structured export format for engines like <strong>Tesseract</strong>.</li>
<li><strong>Cons:</strong> Native support for complex, freeform multi-point baseline curves is limited. While practical for early printed works (incunabula or clean broadsides), hOCR struggles with erratic manuscript layouts and multi-layered marginalia.</li>
</ul>
<h2 id="4-tei-xml-the-academic-and-digital-humanities-benchmark">4. TEI-XML: The Academic and Digital Humanities Benchmark</h2>
<p>The <strong>Text Encoding Initiative (TEI)</strong> format is not strictly an OCR engine output format; rather, it is the premier standard for critical digital editions and scholarly representation of literary and historical texts.</p>
<h3 id="connecting-ocrhtr-to-tei">Connecting OCR/HTR to TEI</h3>
<p>Modern pipelines rarely stop at raw character recognition. Scholars use tools like the <strong>Transkribus TEI Exporter</strong> or automated XSLT pipelines to translate PAGE XML or ALTO files into TEI-conformant XML:</p>
<ul>
<li>Abbreviations are expanded (<code>&lt;choice&gt;&lt;abbr&gt;...&lt;/abbr&gt;&lt;expan&gt;...&lt;/expan&gt;&lt;/choice&gt;</code>).</li>
<li>Deletions, additions, and scribal hands are formally classified (<code>&lt;add&gt;</code>, <code>&lt;del&gt;</code>, <code>&lt;handShift&gt;</code>).</li>
<li>Layout data is preserved alongside literary analysis via the <code>&lt;facsimile&gt;</code> and <code>&lt;surface&gt;</code> elements.</li>
</ul>
<p>If your historical project aims to create an interactive critical edition or semantically searchable scholarly archive, converting your OCR/HTR data into TEI-XML is often the required terminal step.</p>
<h2 id="comparative-matrix-ocrhtr-formats-at-a-glance">Comparative Matrix: OCR/HTR Formats at a Glance</h2>
<table>
<thead>
<tr>
<th style="text-align:left">Feature / Criterion</th>
<th style="text-align:left">PAGE XML</th>
<th style="text-align:left">ALTO XML (v4+)</th>
<th style="text-align:left">hOCR</th>
<th style="text-align:left">TEI-XML</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left"><strong>Primary Domain</strong></td>
<td style="text-align:left">Cursive HTR &amp; Manuscripts</td>
<td style="text-align:left">Mass Library Digitization</td>
<td style="text-align:left">Web OCR &amp; Light Search</td>
<td style="text-align:left">Scholarly Critical Editions</td>
</tr>
<tr>
<td style="text-align:left"><strong>Baseline Support</strong></td>
<td style="text-align:left">Native, Multi-point Polylines</td>
<td style="text-align:left">Supported (since v4.0)</td>
<td style="text-align:left">Basic (Slope/Offset)</td>
<td style="text-align:left">Via facsimile mapping</td>
</tr>
<tr>
<td style="text-align:left"><strong>Irregular Polygons</strong></td>
<td style="text-align:left">Full</td>
<td style="text-align:left">Full</td>
<td style="text-align:left">Limited</td>
<td style="text-align:left">Via coordinate elements</td>
</tr>
<tr>
<td style="text-align:left"><strong>Tooling Ecosystem</strong></td>
<td style="text-align:left">Transkribus, eScriptorium</td>
<td style="text-align:left">METS, Goobi, Kitodo</td>
<td style="text-align:left">Tesseract, Web viewers</td>
<td style="text-align:left">Oxygen, TEI Publisher</td>
</tr>
<tr>
<td style="text-align:left"><strong>Standardization Body</strong></td>
<td style="text-align:left">PRImA Group / Open</td>
<td style="text-align:left">Library of Congress</td>
<td style="text-align:left">Community Specification</td>
<td style="text-align:left">TEI Consortium</td>
</tr>
</tbody>
</table>
<h2 id="practical-recommendations-choosing-your-archival-pipeline">Practical Recommendations: Choosing Your Archival Pipeline</h2>
<p>To establish an efficient, future-proof digitization pipeline:</p>
<ol>
<li><strong>For Pure Handwritten Manuscripts &amp; Archives:</strong><br>
Standardize your transcription and layout extraction on <strong>PAGE XML</strong>. Its baseline calculation and polygon contouring handle non-standard scribal hands with minimal data loss.</li>
<li><strong>For Large-Scale Library &amp; Mixed Collections:</strong><br>
Choose <strong>ALTO XML (v4.2 or higher)</strong> paired with <strong>METS</strong>. This guarantees seamless integration into standard digital repository architectures and digital asset management systems (DAMS).</li>
<li><strong>For Web Presentation &amp; Full-Text Search Indices:</strong><br>
Use <strong>hOCR</strong> or derive lightweight GeoJSON/Web Annotation structures from PAGE XML to drive interactive IIIF viewers (such as Mirador or Universal Viewer) with live in-browser text overlays.</li>
<li><strong>For Scholarly Editions &amp; Paleographic Research:</strong><br>
Generate your ground truth in PAGE XML, perform recognition, and pipe the output through an automated converter to generate <strong>TEI-XML</strong> for editorial markup.</li>
</ol>
<p>By matching the structural capabilities of these formats to the paleographic demands of your source material, you ensure that every stroke, abbreviation, and historical nuance remains decipherable for centuries to come.</p>
<h2 id="frequently-asked-questions-faq">Frequently Asked Questions (FAQ)</h2>
<p><strong>Q1. What is the fundamental difference between standard OCR and HTR?</strong><br>
OCR recognizes consistent machine-printed typography, whereas HTR (Handwritten Text Recognition) utilizes deep neural networks to decode continuous, variable human handwriting and curved baselines.</p>
<p><strong>Q2. Can Tesseract OCR produce PAGE XML or ALTO output for historical documents?</strong><br>
Yes, Tesseract can generate native ALTO XML and hOCR output, and third-party wrappers can convert these results into PAGE XML.</p>
<p><strong>Q3. Why are baselines more important than bounding boxes in handwritten text transcription?</strong><br>
Baselines track the natural, undulating line of human penmanship, allowing software to segregate overlapping ascenders and descenders that collide inside rigid bounding boxes.</p>
<p><strong>Q4. How does the IIIF standard interact with these OCR file formats?</strong><br>
IIIF serves high-resolution images via open web APIs, while formats like ALTO or PAGE XML provide coordinate data that can be converted into IIIF Content Search annotations.</p>
<p><strong>Q5. Which file format is easiest to convert directly into a searchable PDF?</strong><br>
Both hOCR and ALTO XML can be paired with original page images to construct searchable dual-layer PDF files using tools like OCRmyPDF.</p>
<h2 id="see-also">See Also</h2>
<ul>
<li><a href="https://blog.fileformat.com/en/pdf/pdfa-3-the-hybrid-monster-embedding-original-data-inside-your-ocr/">PDF/A-3 - The Hybrid Monster? Embedding Original Data Inside Your OCR</a></li>
<li><a href="https://blog.fileformat.com/ocr/understanding-ocr-file-formats-hocr-vs-alto-vs-pdfa-explained/">Understanding OCR File Formats - HOCR vs ALTO vs PDF/A Explained</a></li>
<li><a href="https://blog.fileformat.com/pdf/what-is-the-difference-between-pdf-and-fdf/">What is the Difference Between PDF and FDF?</a></li>
<li><a href="https://blog.fileformat.com/pdf/what-is-fdf-used-for/">What is FDF Used For? Understanding the Purpose of Forms Data Format</a></li>
<li><a href="https://blog.fileformat.com/file-formats/pdf-vs-word-which-one-should-you-use-and-when/">PDF vs Word: Which One Should You Use and When?</a></li>
</ul>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
