repo-structure: authors/ is canonical cross-section dossier store; section AUTHORS-EN.md now thin indexes

- 56 new author dossiers (all sec04 + sec01 authors), 69 total
- sec04/sec01 AUTHORS-EN.md rewritten as item→dossier link tables
- AGENTS.md conventions + Phase 0 updated
- item 54 identity: Ann Belford Miller (Alice Miller-Ulanov); homonym dossier miller-alice-englard (Alicja Englard, NOT the same person)
This commit is contained in:
Dmitry Kokorin 2026-09-16 10:48:43 +03:00
parent e87656c88d
commit a2527a88ee
118 changed files with 1400 additions and 78 deletions

View file

@ -15,8 +15,9 @@ Sections 1–3, 5–7 are skeletons for later.
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
tools/ # search helpers (lg.py, rsl.py, sx.sh)
authors/ # one dossier per author (created on demand)
authors/ # CANONICAL: one dossier per author (shared across all sections)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
# + AUTHORS-EN.md = THIN INDEX (item → author → link to authors/<slug>.md)
```
## Conventions
@ -31,10 +32,17 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
- **Book file contents:** human-readable editions table + raw **RSL JSON** block +
libgen edition URLs (`https://libgen.vg/edition.php?id=...` — contains download links;
direct file links rotate, don't store them) + notes (identity check, RU↔EN title pairing source).
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) hold: EN identity
(3–5 flagship EN titles, dates), RU surname candidates (all plausible transliterations,
German names often have several: Нойманн/Нойман/Нейманн), verified RU works with EN mapping,
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) are the **CANONICAL** home for
identity, shared across ALL sections (same author can appear in 01, 04, …). One dossier per
PERSON, slug `<lastname>-<firstname>` (lowercase, hyphens). Hold: dates, RU name candidates
(all plausible transliterations, German names often have several: Нойманн/Нойман/Нейманн),
the list items that cite this author (sec # + status), verified RU works with EN mapping,
homonym warnings.
- **`sections/0N/AUTHORS-EN.md` is a THIN INDEX only** — table `item(s) → author → link to
`../../authors/<slug>.md``. It must NOT duplicate identity data (no RU-name analysis, no RU
works table). If an author appears in two sections, the dossier is written once and both
section indexes point to it. Cross-section homonym traps (e.g. the two different "Alice
Miller"s) get their own dossier + a warning in both.
- **Commit per completed book** (or when meaningful author info is gained), message like:
`sec04: <book> — <finding>`.
- Update `CONTEXT.md` for anything useful that fits no category.
@ -256,8 +264,10 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
**Phase 0 — Full names + identity** (`authors/`):
1. Resolve every initial-only author to a full EN name: `tools/ol.py 'title:"EN Title"'`
(author_name) → else `tools/wd.py` (QID + RU label) → else Google Books ISBN page →
else publisher/review page. Record in `sections/<n>/AUTHORS-EN.md` (EN full name, RU name
to query, name source).
else publisher/review page. Write/extend the **dossier** `authors/<slug>.md` (EN full
name, dates, RU name to query, name source); then add a one-line entry in the thin
index `sections/<n>/AUTHORS-EN.md` linking to it. If the dossier already exists from
another section, just append the new list item + link it.
2. For big names: `wd.py` gives life dates + the OFFICIAL RU label (best RSL spelling).
3. Build the transliteration matrix per author (2–3 variants: Нойманн/Нойман, Кох/Коч,
Калфф/Кальфф, Ферс/Фёрт) — run every query with ALL variants.