*A note for historian collaborators on the “Representations of women in Danish West Indies newspapers” project.*
Ordinary keyword search is a poor instrument for this material. A runaway notice that prints *Negro Woman* may be stored as `Ncgro- Womau`. The same woman may appear as *negerinde*, *Negerinden*, *ncgerinde*, or *negeriude*. A three-column gazette mixes Danish and English on one page; Mediestream’s OCR layer — produced by the Royal Danish Library in 2016, not by us — splits words, drops letters, and sometimes reads *Roman* where the type says *Woman*. If we only search for the spelling we have in mind, we systematically miss the people the newspapers describe.
This post explains the retrieval system we have built so that you can use it without installing anything, see how it is constructed, and judge where it helps — and where it still needs you.
—

## 1. What the system is for
The research question is comparative. We want to retrieve, with a stated recall guarantee, every place in the corpus where a woman is *described* (not merely named), so that later estimates of how those descriptions change around emancipation are not artefacts of a leaky search.
The corpus is the official St. Croix gazette, bracketing 1848:
| Period | Years | Title | Issues | Newspaper pages (approx.) |
|---|---|---|---|---|
| Before emancipation (S1) | 1835–1847 | Dansk Vestindisk Regierings Avis, then St. Croix Avis | 1,347 | 5,457 |
| After emancipation (S2) | 1849–1861 | St. Croix Avis | 1,360 | 5,472 |
| Pilot year (outside the comparison) | 1816 | DVRA | 103 | 431 |
1848 itself is excluded: it is the legal break, not a regime. *Sanct Thomæ Tidende* is digitized and would give a St. Thomas contrast; we have not indexed it yet.
The public face of the system is a website. You type a word or phrase; it returns the date, newspaper, page, a snippet of the computer’s reading of the type, and — for passages found by the project’s list of words for women — a crop of the printed notice with the match boxed in red. Each hit links to the corresponding page in Mediestream via a permanent handle (`hdl.handle.net`). Results are listed **oldest first**, because chronology is the natural order of the archive.
—
## 2. How it is built (the pipeline)
The work happens in four stages. Only the last one is the website.
### 2.1 Collect the issues
PDFs come from Det Kgl. Bibliotek Mediestream. Each file is one issue: a modern copyright wrapper, then the scanned pages with an invisible OCR text layer. The newspapers are public domain (Danish print older than 100 years). We did **not** re-OCR the images. The words you will see — including the errors — are the library’s text layer, which is also what a historian using Mediestream’s own search would be matching against.
### 2.2 Cut pages into passages
A whole PDF is too coarse a hit: you want the notice, not “page 2 of 4 September 1843.” Mediestream stores a bounding box for every OCR word. We:
1. Drop the copyright wrapper.
2. Detect columns from the gutters in the word boxes (most pages are three columns).
3. Read each column top to bottom, joining fragments that sit on the same line.
4. Start a new passage when the next line jumps column, moves back up the page, or leaves a vertical gap larger than ordinary leading (on this gazette, notices typically sit 10–25 points apart; lines inside a notice almost overlap).
5. Undo end-of-line hyphenation (`yel-` + `low` → `yellow`).
The unit of retrieval is therefore roughly **one notice or paragraph**, not a page and not a token. The current index has **156,127 passages** across **2,810 issues**. About 4–5% of passages are still too wide (mastheads, tables, badly detected columns). The red box on the crop is still placed from the word coordinates, so even a long passage can point at the right line.
### 2.3 Index every OCR word form
We store passages in SQLite with a full-text index, and we materialize the **vocabulary of the OCR itself**: **673,957 distinct word forms**. That list is the system’s memory of how the type was read — `woman`, `womau`, `ncgro`, `negerinden`, and a great deal of junk. Search does not guess at OCR errors from a dictionary of English or Danish. It looks the query up in *this* vocabulary.
### 2.4 Match with a recall rule, not with “smart” ranking
For each query word of length $\ell$, the system computes an edit threshold $k_s(\ell)$: the smallest number of letter-level edits such that, if the OCR corrupts each character independently with probability at most $\varepsilon_s$, the true printed form is still within $k$ edits with probability at least $1-\eta$.
In the paper this is the binomial quantile of an edit channel (not the looser Hoeffding bound, which would retrieve far too much). Current placeholders, until page audits supply a measured rate, are $\varepsilon = 0.05$ and $\eta = 0.05$. Under those values a ten-letter word is allowed two edits; a four-letter word is matched exactly, because fuzzy `woman` otherwise floods the results with `roman`.
Concretely:
– `negerinde` expands to the attested OCR forms within two edits (`negerinden`, `ncgerinde`, `negeriude`, … — on the order of seventy forms).
– `negro woman` requires both words, in that order, with at most two other tokens between them. It still finds `Ncgro- Womau named KATY` in the 4 September 1843 *Avis*.
– Prefixing a word with `=` forces exact spelling (`=woman` will not also retrieve `roman`).
The expansion is exhaustive over the corpus vocabulary, so the per-word recall statement of the paper carries over. A query of \(m\) words is then guaranteed only by a union bound (\(1 – m\eta\)). Phrases are therefore slightly leakier than single distinctive terms; that is a reason to prefer `negerinde` or `mulatinde` when the Danish form is what you mean.
We do **not** use embeddings, ChatGPT, or “semantic search” at this stage. Those methods have no recall guarantee, and they tend to retrieve passages *about* women rather than passages that *use a descriptor*. The paper’s later estimates need the latter.
—
## 3. Technical notes (for those who want them)
This section can be skipped. It is here so that a computationally inclined colleague can see that the website is the same model as the paper, not a demo with different rules.
**Edit channel.** A printed string \(v\) of length \(\ell\) is modelled as passing through a channel that, independently at each position, copies the character or applies one insertion, deletion, or substitution, with corruption probability at most \(\varepsilon_s\). Then \(\Pr(\mathrm{ed}(v,\tilde v)\le k)\ge F_{\ell,\varepsilon}(k)\), the binomial cdf. We set \(k_s(\ell)=\min\{k:F_{\ell,\varepsilon_s}(k)\ge 1-\eta\}\).
**Recall.** If at least a fraction \(\rho_s\) of true mentions have their *printed* (not OCR’d) anchor in the lexicon, expected recall is at least \(\rho_s(1-\eta)\). \(\rho_s\) is a historical-linguistic quantity. It cannot be estimated from OCR text; it needs the page-image audit (H2).
**Implementation.** Offline, RapidFuzz compares each query word to vocabulary strings of nearby length; a \(q\)-gram filter discards hopeless candidates before the Levenshtein distance is computed. Online, the GitHub Pages site runs the *same* expansion and proximity rule in the browser over a static export of the index (vocabulary files by word length, postings sharded by a hash of the form, passage text in shards of 1,000). No search server is required. A typical distinctive query returns in under a second; the whole women lexicon takes a few seconds.
**Scoring** is used only to break ties and for the optional “best match first” view: quality \(1 – d/\ell\) times an IDF term so that rare OCR forms rank above ubiquitous ones. Default display order is chronological (then page, then passage id).
**What still leaks.** Tokenisation cannot recover a word the OCR split (`wo` + `man`) or glued to its neighbour. Passages wider than half a page are a layout failure, not a matching failure. Repeated advertisements across successive issues are retrieved as separate hits: passage counts are not counts of distinct women.
—
## 4. How to use the website
You need a browser and the site address (once it is published on GitHub Pages). Nothing to install.
**Search.** Type a word or a short phrase and press Search. Several words must appear close together and in that order: `negro woman` finds “a Negro Woman named Jane,” not a page that contains *negro* in one column and *woman* in another. Buttons under the box run examples; **All words for women** runs the project lexicon in one go.
**Read a hit.** Each card shows the date, title, newspaper page, and whereabouts on the page (“right column, middle”). Yellow highlighting is the computer’s reading. The crop, when present, is the scan. **Always trust the picture.** The Royal Library link opens that page in Mediestream (2,804 of 2,810 issues have a verified page identifier; the remaining six have no link rather than a wrong one).
**Filters.** Restrict to S1 / S2, a year range, or one title. **Exact spelling only** turns off OCR expansion for every word in the query.
**Sort.** Oldest first (default), newest first, or best match first. Changing sort does not re-run the search.
**Spellings included in this search.** Open this panel to see which OCR forms were admitted. Untick anything that is clearly the wrong word (`roman` for `woman` is the usual example), then **Search again**. That panel *is* the first round of lexicon validation (H1).
**Checks.** For each passage: **Yes** if it describes a woman, **No** if it does not, **Unsure** otherwise. Optional note. Put your name in the bar at the bottom, then **Download my checks**. The file is a spreadsheet (date, newspaper, page, your label, the snippet, the Mediestream link). Checks live only in your browser until you download them; nothing is uploaded.
**Share a search.** **Copy link to this search** puts the query, filters, and sort into the address. You can email a colleague “please look at this.”
**Pictures.** Crops are pre-rendered for passages the women lexicon retrieves (about 2,710 notices). Other queries still return text and the library link; we cannot crop all 156,000 passages into a site that GitHub Pages will host, and Mediestream’s image service cannot be called live from the page.
A worked example: search `negro woman`, period “Before emancipation.” You should see, among others, the 4 September 1843 sale of “a young and healthy Negro-Woman named KATY,” and the May 1835 runaway notice for Jane. The OCR says `Ncgro- Womau` and `Xegro Woman`. The red box sits on the printed words.
—
## 5. How this fits the historian tasks
The website is a working instrument, not a substitute for the four inputs in the project plan.
| Input | What the site already does | What we still need from you |
| — | — | — |
| **H1** Seed terms and variants | Machine-generated OCR neighbours, shown in “Spellings included…” | Accept / reject / add forms; flag terms that shift meaning around 1848; flag terms too ambiguous to search alone (`girl`, `pige`) |
| **H2** Page audits | A frozen random sample of 40 pages per stratum already exists (seed 20260922). Do not pick extra pages | Read the *image*, transcribe every description of a woman, copy two or three lines for the character-error rate |
| **H3** Passage labels | The lexicon search is the candidate pool | Labels on a **random** sample of hits, in two batches — not only the notices you chose to click |
| **H4** What the source can tell us | Nothing automatic | A short memo: which facets refer to a fact independent of the newspaper’s classification, and how often the paper is wrong |
A caution about H3. Browsing the site and marking Yes/No is invaluable for cleaning the lexicon and for getting a feel for the genre. It is **not** the labelled sample the estimates require, because you will have chosen interesting hits. The paper’s difference estimator is unbiased only if the labelled set is a probability sample of retrieved passages. We will draw that sample from the index and send it as a list; the site can still be used to look each row up.
—
## 6. Can this be used on other historical collections?
Yes, with the same architecture, provided four conditions hold.
**You need page images with word-level coordinates**, or a way to get them. Mediestream PDFs already contain an OCR layer with boxes; Chronicling America ALTO XML, Transkribus PAGE XML, and many library IIIF services do too. If you have only plain text, we can still retrieve, but we cannot crop a notice or point at a column. If you have only images, we would run a modern OCR (that is a different \(\varepsilon_s\), and the recall numbers have to be re-estimated).
**The unit of interest should be a passage, not a volume.** The column-and-gap splitter is tuned to this gazette’s leading and three-column layout. A book, a court protocol, or a Fraktur *Intelligenzblatt* would need the splitter re-tuned — or replaced by layout models trained on that genre. The matching rule does not care; the geometry does.
**Someone who knows the language of the source must own the lexicon.** The machine proposes OCR neighbours. It does not know that *frimulatinde* is a legal status in this colony and a slur in another, or that *wench* in a St. Croix sale notice is not the same word as in a London farce. A new project begins with seed terms for *that* archive, in *those* languages, and a historian who will sit with the variant list.
**The recall guarantee is only as good as \(\varepsilon_s\) and \(\rho_s\).** Both come from a random page audit of the new collection, not from our 1816 pilot. Fraktur, faint microfilm, and mixed typefaces will have a higher \(\varepsilon_s\); the thresholds \(k_s(\ell)\) will open automatically, and the historian will spend more time rejecting spurious variants. That is the system working as designed, not failing.
What does **not** transfer without thought:
– The women lexicon (it is specific to this bilingual colonial gazette).
– The S1/S2 cut at 1848 (that is a historical argument, not a software setting).
– The assumption that “one notice ≈ one passage” (it fails on running text and on tables).
– Any claim about people rather than about print. Retrieval finds descriptions. Whether a description refers to a fact is H4, and it is not a retrieval problem.
Projects that fit well: runaway and sale notices in other Caribbean or North American gazettes; formulaic legal announcements; parish registers with a stable vocabulary and a known OCR layer. Projects that fit badly, unless we change the model: open-ended “find me everything about X,” where X has no lexical anchor.
—
## 7. What the system will not do
It will not tell you whether KATY was twenty-seven, a good cook, or enslaved in the sense a plantation register would record. It will not de-duplicate the same advertisement run three times. It will not read a name with no descriptor as a description of a woman. It will not protect you from the fact that these pages name people who were bought, sold, and hunted in print: the site is public because the scans are public, but how we quote, anonymise, and consult is an ethics question for the group, not a search setting.
The honest use of the system is narrower and, we think, more useful. It makes a noisy official gazette *searchable as print*, with a rule you can argue with, a picture you can check, and a list of spellings you can veto. Everything that turns those hits into historical claims still runs through your reading.
—
## 8. A first hour on the site
If you have an hour:
1. Search `negerinde`, then `negro woman`, then `free coloured`. Note which genres appear (sale, runaway, mortgage, news filler).
2. Open **Spellings included in this search** for `woman`. Untick the forms that are not *woman*. Search again.
3. Switch the period filter from All years to Before / After emancipation and look at the year bars — not as a result, only as a hint of volume.
4. For ten hits, mark Yes / No / Unsure against the **image**. Download the spreadsheet.
5. Send us: (a) that file; (b) any seed term we are missing; (c) any term in the lexicon that should never be searched in isolation.
That is already Task 1, a sketch of Task 5, and a feel for why Task 2 has to be a random page rather than the notices we all find first.
—
*Newspaper scans and OCR: Det Kgl. Bibliotek, Mediestream. Index and website: this project. The lexicon is a placeholder until historians validate it. The OCR error rate used for the edit threshold is a placeholder until the page audit.*
Leave a comment