About the archive
What is here
The archive holds 627 works — 18,709 files, 79 GB on disk. The oldest carries a date of 1827, the newest 2026. An encyclopedia of 21,546 articles has been written from them, 898 of those short sourced notices grouped by year, and each one names the sources it used.
The two are counted separately on purpose. A work can sit here unread: 508 of the 627 are not cited by any article yet. The bibliography lists those anyway, each with a provenance record saying where it came from and when.
What was collected
By class of source, largest first — the file count is what the archive holds, not what the institution has.
- Newspapers & periodicals
- 17 works · 9,638 files · 54 GB
- Websites archived
- 3 works · 4,216 files · 2.1 GB
- Government & municipal records
- 9 works · 1,592 files · 873 MB
- Newspapers — selected pages only
- 431 works · 1,305 files · 42 MB
- Maps & charts
- 15 works · 663 files · 11 GB
- Audio recordings
- 1 work · 428 files · 3.1 GB
- Manuscript & archival transcriptions
- 4 works · 309 files · 23 MB
- Books & serials
- 131 works · 286 files · 3.3 GB
- Geospatial & scientific layers
- 14 works · 248 files · 4.9 GB
- Photographs & page images
- 2 works · 24 files · 346 MB
The deepest run is the Beaver Beacon, 626 issues from 1955 to 2011. Behind it sits the Charlevoix County Herald, 2,644 issues from 1902 to 1953 — the county paper, covering the island from Charlevoix. Almost none of it was digitised here. The scans came chiefly from Internet Archive, from Charlevoix Public Library microfilm; Beaver Island District Library / Beaver Beacon; and Library of Congress, Chronicling America. The credits name every institution behind every collection.
Pictures, maps and sound
6,373 objects inside those collections are works in their own right and have a page each — a title, the rights exactly as the holder stated them, and the file itself. Browse them at the media index, or in date order through the gallery.
The split that matters is between what was collected and what was drawn here. 6,303 are originals: someone else photographed, surveyed, printed or recorded the thing, and this archive holds a copy under whatever terms they set.
- Photographs
- 5,869
- Nautical charts
- 186
- Page extracts
- 90
- Maps
- 78
- Sound recordings
- 70
- Plates — this archive's own cartography
- 5
- Moving images
- 5
The photographs are recent, and almost all of them are other people's. The largest groups, named as the holders name them: Community photographs — public Facebook groups and pages (2,859); Beaver Beacon photo gallery (1,886); beaverisland.net community site (737); Beaver Island Historical Society website (264). Each image carries the credit its poster or publisher gave it, and the rights page sets out what may be shown and what may not.
The charts are NOAA's historical collection — 186 survey sheets of the archipelago and its approaches, 1855 to 2024, scanned from the originals. The 78 maps come from the USGS, the General Land Office's survey plats, and a handful of libraries holding older sheets.
All the sound is one collection. Alan Lomax, Michigan folk-song recordings, Beaver Island (1938) runs to 70 items — 6 hours 16 minutes of singing and talk, recorded for the Library of Congress. Each has a transcript, so a reader who cannot play the audio can still read what was sung.
The other 70 were made here.
- Plates — this archive's own cartography
- 70
They are drawings, rendered in Python: matplotlib figures written by 13 scripts in the data repo, projected to Michigan GeoRef (EPSG:3078) from the geospatial layers listed further up this page. Every plate's page names the script that drew it and the data it was drawn from. Nothing in the atlas is a scan of a historical map, and nothing listed above it was drawn here. The plates are licensed CC BY 4.0; the data underneath keeps its own terms and is credited separately.
What the search covers
The search page has two tabs. One searches the encyclopedia. The other searches the OCR text of the sources themselves — 521,212 segments drawn from 2,778 transcribed files, across five collections.
- Charlevoix newspapers
- 450,608 segments from 948 files
- Books
- 57,169 segments from 128 files
- Beaver Beacon
- 11,918 segments from 625 files
- Island websites
- 983 segments from 983 files
- Archival transcriptions
- 534 segments from 94 files
A hit is a segment, not a page. The text is split on paragraph boundaries near 3 kB and no page markers survive it, so a result points at a source file and the path printed under it is that segment's provenance. Fuzzy matching is on by default because the text is machine-read: Gallaher finds Gallagher. Most of the newspaper text comes from bound volumes carrying no per-issue date, so the year filter reaches a minority of segments and excludes the rest rather than guessing at them.
Transcription is still running, so these figures rise. What a machine transcription can and cannot be trusted to say — including the failure mode where it invents fluent prose for a page it could not read — is set out in the editorial standards.
Every figure above is read at build time from an export in the data repo: works, files, bytes and citation counts from the source catalogue; media counts and durations from the media manifest; article counts from the knowledge base itself. Segment counts come from the search index manifest, written when the index was last built on 4 August 2026, and describe what was loaded into it — not a live query.