8th International Workshop on Historical Document Imaging and Processing (HIP’26)

3-4 September 2026, Vienna, Austria

There have been increased efforts worldwide to digitize our cultural heritage conveyed in historical documents. In this workshop, we bring together researchers from various fields working on document image acquisition, restoration, analysis, indexing, and retrieval to make these documents accessible in digital libraries.

It is the eigth satellite workshop of ICDAR dedicated to this topic, following HIP’11 in Beijing, HIP’13 in Washington, HIP’15 in Nancy, HIP’17 in Kyoto and HIP’19 in Sydney, HIP’21 in Lausanne (hybrid) and HIP’23 in San José, that were a significant success with strong participation.

The workshop is planned for 1½-days with oral presentations on September 3rd, and an excursion (to be confirmed) on September 4th. Each submission will undergo peer-review and distinguished submissions will be presented orally.

HIP aims to provide the researchers with a forum that is complementary and synergetic to the main sessions at ICDAR on document analysis and recognition. The manifold topics addressed in this workshop encompass the entire processing chain from image acquisition to information extraction. We include the growing importance of machine learning in this processing chain, and we encourage the presentation of entire projects in the context of historical documents.

UPDATE:
The workshop program with accepted papers and session assignments is now published.
Each accepted paper has been allocated 15min for presentation + 5min for questions.

UPDATE:
A brief extension to the submission deadline will be granted.
The new deadline is 26 May 2026 AoE.

Submission Deadline:                 22 May 2026 (Time zone: Anywhere on Earth)
                                                        26 May 2026 (Time zone: Anywhere on Earth)

Acceptance Notification:            22 June 2026

Camera Ready:                             7 July 2026

Workshop:                                    3 September 2026

Excursion:                                     4 September 2026 (to be confirmed)

Venue:                                           TU Wien (Campus Gußhaus – Faculty of Electrical Engineering and IT, Gußhausstraße 27-29, 1040 Vienna)

Call for Papers

It is our pleasure to announce that the 8th International Workshop on Historical Document Imaging and Processing (HIP’26) will be held in conjunction with ICDAR2026, on 3-4 September 2026 in Vienna, Austria.

The workshop brings together researchers working with historical documents and intends to be complementary and synergistic to the work in analysis and recognition featured in the main sessions of ICDAR, the premier international forum for researchers and practitioners in the document analysis community.

Submissions are received until 22 May 2026 26 May 2026 (Time zone: Anywhere on Earth) via CMT and undergo review by the members of the Program Committee.

Submissions must follow the ICDAR guidelines and template provided. It is not required to anonymize the submission, but authors are welcome to do so if they prefer it. Acceptance notifications will be sent out 22 June 2026, with camera ready submissions due on 7 July 2026.

Workshop topics include (but are not limited to):

Imaging and Image Acquisition

  •  Imaging for fragile materials
  •  Multispectral imaging
  •  Camera-based/non-invasive acquisition
  •  Case studies/applications

Digital Archiving Considerations

  • Compression issues
  • Measuring essential resolution (colour, spatial) and metadata
  • Modelling of document image degradation
  • Historical Collections
  • Military records, personal journals, church records, medieval manuscripts, etc.
  • Scientific, technical and educational documents
  • Government archives, documents from the world cultural heritage, multi-language

Document Restoration/Improving readability

  • Removing or minimizing damages, defects, ink-bleed
  • Completing and filling in missing pieces based on context, prior knowledge
  • Machine-learning algorithms for enhancement based on example images
  • Interactive tools from a user viewpoint
  • Learning from user-directed image enhancement

Document Content Acquisition and Information Extraction (within the context of historical documents)

  • Automated or semi-automated transcription (OCR, OLR, HTR)
  • Machine-learning algorithms for content extraction, including recurrent neural networks, auto-encoders, transformers, and unsupervised feature learning
  • Content recognition based on surrounding and supporting context
  • Annotation
  • Evaluation metrics and methods
  • Ontologies for modelling historical document content
  • Content-based retrieval

Family History Documents and Genealogies

  • Personal, Family, National and Historical Collections of Family Genealogy and Histories
  • Extracting and linking names, dates, places, etc.
  • Extracting, linking and piecing together personal and family histories and narratives
  • Discovering historical social networks

Automated Classification, Grouping and Hyperlinking of Historical Documents

  • Style identification (of printed text/handwriting, dating or author identification)
  • Searching for Documents over the Internet
  • Web-based navigation within/among document images
  • Search/query, retrieval, summarization or condensation of document images
  • Document collecting, clustering, linking and analysis technologies
  • Parallel tagging of images, transcripts, and other document layers

Digital Humanities applications of document analysis and recognition

  •  Computer Vision for Computational History
  •  Digital methods and tools for the study of historical documents
  •  Crowdsourcing

Artificial Intelligence and Machine Learning for historical documents

  • VLMs and LLMs for analysis and recognition of historical source materials
  • Training, fine-tuning and evaluation of specialized AI models
  • Prompt strategies for historical data and contexts

For work focusing on handwriting/paleography, we recommend you have a look at IWCP: 4th International Workshop on Computational Paleography.

The Microsoft CMT service was used for managing the peer-reviewing process for this conference. This service was provided for free by Microsoft and they bore all expenses, including costs for Azure cloud services as well as for software development and support.

General Chair

Clemens Neudecker
Berlin State Library
Directorate General
Potsdamer Strasse 33
10785 Berlin
Germany
clemens.neudecker@sbb.spk-berlin.de  

HIP Series Chair

Apostolos Antonacopoulos
PRImA Research Lab
School of Science, Engineering & Environment
University of Salford
Greater Manchester M5 4WT
United Kingdom
a.antonacopoulos@primaresearch.org

Program Chairs

Maud Ehrmann
EPFL CDH DHI DHLAB
INN 116 (Bâtiment INN)
Station 14
CH-1015 Lausanne
Switzerland
maud.ehrmann@epfl.ch

Christian Clausner
PRImA Research Lab
School of Science, Engineering & Environment
Newton Building
University of Salford
Greater Manchester M5 4WT
United Kingdom
c.clausner@primaresearch.org

Publications Chair

Kai Labusch
Berlin State Library
Information and Data Management
Potsdamer Strasse 33
10785 Berlin
Germany
kai.labusch@sbb.spk-berlin.de

Local Arrangements Chair

TBD

Honorary Chair

William Barrett
Department of Computer Science
Brigham Young University
Provo, Utah 84604
USA
barrett@cs.byu.edu

Program Committee
  • Andreas Fischer, Switzerland
  • Anna Scius-Bertrand, Switzerland
  • Basilis Gatos, Greece
  • Chahan Vidal-Gorène, France
  • Christopher Kermorvant, France
  • Daniel Stoekl, France
  • David Doermann, USA
  • Dominique Stutzmann, France
  • Elisa Barney Smith, USA
  • Irina Rabaev, Israel
  • Isabel Marthot-Santaniello, Switzerland
  • Josep Llados, Spain
  • Katherine McDonough, United Kingdom
  • Kengo Terasawa, Japan
  • Lars Vögtlin, Switzerland
  • Marcel Bollmann, Sweden
  • Maroua Mehri, Tunisia
  • Masaki Nakagawa, Japan
  • Mickael Coustaty, France
  • Nicholas Howe, USA
  • Peter Stokes, France
  • Philip Ströbel, Switzerland
  • Rafael Dueire Lins, Brazil
  • Roger Easton, USA
  • Simone Marinai, Italy
  • Solène Tarride, France
  • Tan Lu, Belgium
  • Thibault Clérice, France
  • Thierry Paquet, France
  • Tobias Hodel, Switzerland
  • Tomo Miyazaki, Japan
  • Tracy Powell, New Zealand
  • Verónica Romero, Spain
  • Vincent Christlein, Germany
  • Volodymyr Rybkin, Russia
  • William Barrett, USA

3 September 2026

09:00 – 17:00 | Workshop

09:00-09:10 Welcome message
SESSION 1: Handwritten Text Recognition Models & Transfer (Chair: N.N.)
09:10-09:30 MEDUSA: A Curriculum-Based Vision-Language Framework for Multilingual Medieval Handwritten Text Recognition Théo Moins, Brenna Hensley, Florian Cafiero, Jean-Baptiste Camps, Lilla Conte, Emilie Guidi, Katarzyna Kapitan, Carolina Macedo, Cecile Vermaas, Chahan Vidal-Gorène
09:30-09:50 Language Similarity and Cross-Lingual Transfer in Historical HTR: Evidence from Swedish, Norwegian, and Medieval Latin Micaella Bruton, Crina Tudor, Wout Sinnaeve, Oreen Yousuf, Signe Rirdance, Raphaela Heil, Beáta Megyesi
09:50-10:10 Is a Generic Dataset and Foundation VLM for Arabic HTR Worth It? Lessons from AMIDDA Chahan Vidal-Gorène, Noëmie Lucas, Clément Salah, Aliénor Decours-Perez
10:10-10:30 Beyond Monolithic OCR: Mixture-of-Experts Specialisation for Ancient Devanagari Script Recognition Vriti Sharma, Rajat Verma, Rohit Saluja
10:30-11:00 Coffee Break
SESSION 2: Document Structure & Segmentation (Chair: N.N.)
11:00-11:20 DIVA-UTP: A Pixel-Level Region Labeling Dataset for Complex Layout Analysis of Medieval Manuscripts Najoua Rahal, Rolf Ingold
11:20-11:40 A Diachronic Multi-Type Dataset for Document Layout Analysis from the 17th Century to the Present Juliette Janès, Sarah Bénière, Benjamin Kiessling, Lucence Ing, Eric Astier, Simon Gabay, Benoît Sagot, Thibault Clérice
11:40-12:00 Generalization of Text Line Segmentation for HTR in Historical Documents Gayan Pathirage, Stephan Unter, Simon Corbillé, Elisa Barney Smith
12:00-12:20 Consistent Line Extraction on Medieval Charters: Mask R-CNN and U-Net Architectures
Nicolas Renet, Anguelos Nicolaou, Georg Vogeler
12h20-12h30 Discussion
12h30-13h30 Lunch Break, Campus Gusshaus
SESSION 3: Specialised Text Recognition & Retrieval (Chair: N.N.)
13:30-13:50 Postcorrecting OCR with LLMs: a low resource approach Valentina Vavassori, Harry Lloyd
13:50-14:10 Reconstruction Error Ratios for Prototype-Anchored Unsupervised Learning in Optical Character Recognition Tim Hallyburton, Anna Scius-Bertrand, Arthur Neto, Andreas Fischer, Gernot Fink
14:10-14:30 Cross-View Retrieval of Byzantine Monograms with Contrastive Representation Learning Saranga Mahanta, Victoria Eyharabide, Béatrice Caseau, Isabelle Bloch
14:30-14:50 Diagram Recognition for Byzantine Manuscripts Using a Pretrained nnU-Net Lilly Osburg, Germaine Götzelmann, Felix Kraus, Danah Tonne
14:50-15:00 Discussion
15:00-15:30 Coffee Break
SESSION 4: Newspaper Analysis & Historical Information Extraction (Chair: N.N.)
15:30-15:50 Reading Order Article Coherence Dataset and Measure for Quality Assessment in Historical Newspapers Kai Labusch, Konstantin Baierer, Michał Bubula, Jana Götze, Jörg Lehmann, Clemens Neudecker, Vahid Rezanezhad
15:50-16:10 Towards Hierarchical Structure Understanding of Newspaper Images William Mocaër, Solène Tarride, Thomas Constum, Merveilles Agbeti-Messan, Tom Simon, Clément Chatelain, Stéphane Nicolas, Pierrick Tranouez, Sébastien Cretin, Thierry Paquet
16:10-16:30 Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press Chahan Vidal-Gorène, Seda Kirakosyan, Edita Matevosyan
16:30-16:50 A Prompt Optimization Framework for Parsing Historical Addresses Amel Gader, Rafael Patronilo, Mahsa Vafaie, Genet-Asefa Gesese, Harald Sack
16:50-17:00 Discussion and wrap-up

4 September 2026

Workshop excursion (to be confirmed).

For general enquiries about HIP’26 please contact the organizers.

Copyright © HIP’26 Organizing Committee

Web hosting provided by the Berlin State Library.