jung/AGENTS.md

417 lines
32 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Jung Reading List → Russian Editions: Research Notes
## Goal
For every book in the ISAP Zurich reading list (https://isapzurich.com/en/library/reading-list),
find the **Russian edition(s)**: correct Russian title, publisher, year, translator, ISBN,
libgen download page (if exists), and РГБ (RSL) bibliographic record (if exists).
We work **section by section**. Current: **Section 04 — Pictures** (`sections/04-pictures/`).
Sections 1–3, 5–7 are skeletons for later.
## Repository layout
```
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
tools/ # search helpers (lg.py, rsl.py, sx.sh)
authors/ # CANONICAL: one dossier per author (shared across all sections)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
# + AUTHORS-EN.md = THIN INDEX (item → author → link to authors/<slug>.md)
```
## Conventions
- **One file per book**, named `NN-<slug>.md` (NN = order in the reading list).
- **Status marker** at top of each book file:
- `⬜ not searched yet`
- `🔎 searching in progress`
- `✅ RU edition(s) found`
- `🔶 partial` (e.g. only an essay/chapter available in RU)
- `❌ no RU edition found (verified)`
- **Book file contents:** human-readable editions table + raw **RSL JSON** block +
libgen edition URLs (`https://libgen.vg/edition.php?id=...` — contains download links;
direct file links rotate, don't store them) + notes (identity check, RU↔EN title pairing source).
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) are the **CANONICAL** home for
identity, shared across ALL sections (same author can appear in 01, 04, …). One dossier per
PERSON, slug `<lastname>-<firstname>` (lowercase, hyphens). Hold: dates, RU name candidates
(all plausible transliterations, German names often have several: Нойманн/Нойман/Нейманн),
the list items that cite this author (sec # + status), verified RU works with EN mapping,
homonym warnings.
- **`sections/0N/AUTHORS-EN.md` is a THIN INDEX only** — table `item(s) → author → link to
`../../authors/<slug>.md``. It must NOT duplicate identity data (no RU-name analysis, no RU
works table). If an author appears in two sections, the dossier is written once and both
section indexes point to it. Cross-section homonym traps (e.g. the two different "Alice
Miller"s) get their own dossier + a warning in both.
- **Commit per completed book** (or when meaningful author info is gained), message like:
`sec04: <book> — <finding>`.
- Update `CONTEXT.md` for anything useful that fits no category.
## Sources and their status (verified 2026)
### 1. РГБ (Russian State Library) via libgen.vg biblioservice — PRIMARY for existence + editions
- Tool: `tools/rsl.py "Full Name"` (JSON, polite, 3s between pages)
- Raw URL: `https://libgen.vg/biblioservice.php?value=<query>&type=rsl&format=json`
- Rich records: title parts, **author with life dates** `[1905-1960]` (great identity anchor),
translator, publisher, city, year, pages, ISBN, series, UDC tags, even TOC field.
- **Limitations:** no pagination (fixed first page) → always query by **full first+last name**
for people, or 2–3 distinctive words. Loose matching (searches author AND title fields).
- Other `type=` values on this mirror are **dead**: worldcat, googlebooks, isbndb, udc
(all return "Nothing found" even for control queries). `type=isbn` decodes ISBN structure:
checksum error + **registered publisher group name** (e.g. 5-519 → «Клуб Касталия»,
5-98712 → «Петроглиф») — cheap, unlimited; used by `tools/isbnval.py` as the structure layer.
### 2. libgen.vg full-text search — PRIMARY for readable editions
- Tool: `tools/lg.py "query"` (whole words only, res=100, prints pagination hint)
- Raw URL: `https://libgen.vg/index.php?req=<query>&res=100`
- **Requires a browser User-Agent** (otherwise nginx default page / robot block).
- **Search semantics (verified):** whole-word, case-insensitive. Stems do NOT work reliably
(`происхожден` → 0 hits; full `Происхождение` → many). Multi-word = AND.
→ Use **full inflected words** (nominative, as stored in titles) and 2–3 distinctive words.
- Author search works (`req=Нойманн`) but pulls homonyms — filter by first name in results.
- Pagination: `&page=N` works.
- Edition page: `https://libgen.vg/edition.php?id=<id>`.
### 3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
- URL: `https://cogito-shop.com/search/?q=<query>` (param is **q**, not query)
- Matches author AND title (fuzzy — «исцеляющее сновидение» missed, «Майер» hit). Good for
"is it sellable now in RU".
- **Product pages work with plain curl** (2026-09-19): «Характеристики» block in HTML —
Автор / Издательство / Формат / Вес / Тип обложки / Кол-во стр / Год / ISBN / Код. Search URL
= `/search/?q=<author>`, product link = `/catalog/<cat>/<slug>/`.
### 4. SearXNG (local) — cross-verification
- Tool: `tools/sx.sh "query"` (JSON API at http://localhost:8888; the web_search tool blocks localhost)
- Use to verify RU↔EN title pairing in a third party's words, resolve transliterations.
- Keep queries ≤ 3–4 words; «…» quotes only on distinctive RU titles. Unstable relevance —
drop junk queries, don't retry blindly.
- **Engine status (2026-09-19):** container `searxng` (docker, config
`~/.searxng/settings.yml` → mounted /etc/searxng/settings.yml). Currently **google only**:
bing is BLOCKED (responds but ignores query — random promo junk), brave dead (403),
duckduckgo CAPTCHA, wikipedia/wikidata dead/timeout — all disabled in config.
Diagnosis: `docker logs searxng | grep -oE '(ERROR|WARNING):searx.(engines|network).[a-z_]+' | sort | uniq -c`
+ per-engine test `curl --get :8888/search --data-urlencode 'q=...' --data-urlencode format=json --data-urlencode 'engines=google'`.
Fix = mark engine `disabled: true` in settings.yml + `docker restart searxng`.
- **Always pass explicit `language=ru` / `language=en`** (user-confirmed fix; per-request, config stays multilingual).
### 5. ru.wikipedia / en.wikipedia — author dossiers
- ru-wiki articles often have «…на русском» sections listing actual RU titles.
- en-wiki for the EN bibliography (identity anchor: 3–5 flagship titles).
### 6. livelib.ru author pages — all RU works of a known author
- e.g. `livelib.ru/author/<id>-<slug>` lists every RU edition (server-rendered, scrapable).
- Author ID: find via SearXNG `site:livelib.ru <author>`. The /search page is JS-only (nope).
- **Publisher pages** `/publisher/<id>-<slug>` = the publisher's full book list (server-rendered,
book links `/book/<id>-<slug>` carry title + author slugs) — cheap reverse publisher sweep.
Example: /publisher/2076-medkov-s-b (Медков С. Б. — small esoteric press w/ Jung line).
### 7. cogito-shop person pages — shop stock per author
- `cogito-shop.com/person/<first>_<last>/` (transliterated, first-name first,
e.g. `/person/mariya_luiza_fon_frants/`, `/person/yaffe_aniela/`).
- Slug discoverable from the shop's search page (results contain person links) or guessable.
- Lists every edition the shop carries (title variants!). Product pages lack full
metadata in HTML (JS-rendered) — use RSL/livelib for publisher/year/ISBN.
### 8. GNB SPb (State Public Library of St. Petersburg) catalog — SECONDARY bibliographic source
- URL: `https://www.gbs.spb.ru/ru/search/detail/?id=<hash>` (rich GOST records: title,
responsible parties incl. translators, publisher w/ full legal names, year, pages, ISBN,
notes like «Др. кн. авт.»)
- **Complements RSL:** RSL (RGB) misses some small-press editions (e.g. Pattis-Zhoya
«Аборты…» Т8/ЦГИ 2017 is in GBS but NOT in RSL). Query via SearXNG `site:gbs.spb.ru <author>`.
### 9. chitai-gorod.ru / ozon / labirint — trade retail pages
- chitai-gorod author pages list RU works (some noise). **Product pages: CRW works** (2026-09-19) —
«Характеристики» block: Год издания / Кол-во стр / Переводчик / Издательство / ISBN. Plain curl
gives nothing (JS). Find product IDs via SearXNG `site:chitai-gorod.ru "<RU title>"`.
- labirint.ru book pages work with plain curl — «Характеристики» block after <h1> (publisher,
year, pages, translator, ISBN). Labirint SEARCH is anti-bot (text= ignored).
- Ozon product pages: blocked from curl (redirect loop) AND via CRW (FAB challenge 2026-09-19);
fetch_content also fails. Use SearXNG OZON snippets instead.
- vse-svobodny.com — works with `Cookie: beget=begetok` (anti-bot challenge, 2026-09-19);
product pages carry publisher/year/pages.
#### 10. flibusta — big e-book library. .su = server-rendered search; **.is = OPDS + direct downloads (user fixed access, 2026-07-16)**
- **.is OPDS (preferred):** Tool `tools/flib.py`:
- `flib.py authors "query"` → `/opds/search?searchType=authors&searchTerm=…` (id, name, book count)
- `flib.py books "query"` → `searchType=books` (id, title, author, year, format, translator from annotation)
- `flib.py author <id>` → `/opds/author/<id>/alphabet` — book list (**20 per page!**)
- `flib.py authorall <id>` — follows `rel="next"` pages (`/alphabet/1`, `/2`, …) to the end — **always use authorall** (Jung a/5272 = 128 books; the bare feed shows only the first 20)
- Author HTML: `/a/<id>` (e.g. a/57639 = фон Франц, a/118921 = Нойманн); book page `/b/<id>`
- **Direct downloads:** `/b/<id>/fb2|epub|mobi|html|txt` (pdf/doc as «/b/<id>/download» links on the pages)
- **Quirks (2026-07-17):** `fb2+zip` format actually returns a ZIP (unpack: 1 fb2 + cover inside); `/b/<id>/pdf` sometimes 302-redirects to the book HTML page (a ~20 KB stub) when the file isn't served directly — check downloaded «pdf»s with `file`, fall back to libgen (same editions usually there) or the fb2 twin.
- **Converter quirk (2026-07-18):** `/b/<id>/epub` (or any fmt endpoint) on a book whose NATIVE format is pdf/djvu/doc may return the **native file, not an epub** — always magic-byte-check downloads (`file`); some «20 KB HTML stubs» are in fact real native files with the wrong extension.
- .su = **mirror of .is with reduced functionality** (user, 2026-07-17) — use .is OPDS only.
.su `booksearch/?ask=` exists but adds nothing; API endpoints blocked (api/search.php → 403).
- Value: (1) RU full-text downloads for the download phase; (2) author-book-list = cheap author sweep;
(3) negative verification (all same-name authors listed). Verified 2026-07-09 (.su): no «Выявляющий
образ», no «Сэндплей» titles, no «Тест Дерево», no «Искусство и творческое бессознательное» (as titled).
Flibusta search is fuzzy (whole-phrase not required).
### 10b. OpenAlex + Crossref (`tools/oa.py`) — Phase 0 identity/ISBN (free, no key, better than OL for modern academic books)
- `oa.py "Exact EN Title"` (OpenAlex works: full author names, year, DOI, publisher),
`oa.py --crossref "Title words"` (Crossref: **ISBN list** + publisher + year),
`oa.py --author "Last, First"` (OpenAlex author entity), `oa.py both "Title"`.
- Verified 2026-07-17: Evers-Fahey full name "Karen Evers-Fahey" + Routledge 2016;
"Coming into Mind" = **Margaret** Wilkinson (not Michael!) 9781317710578;
"The Symbolic Quest" Whitmont 9780691213187 (Princeton).
- Use in Phase 0 INSTEAD of / alongside OL; OL still primary for pre-1990 titles.
### 11. Open Library — the "English RSL" (author works lists, full names, ISBNs)
- Tool: `tools/ol.py "author:Lastname, First"` or `tools/ol.py 'title:"Exact EN Title"'`
- `openlibrary.org/search.json?q=...&fields=title,author_name,publish_year,language,publisher,isbn`
- No key, no rate limits observed (1s between calls is polite). Aggregates all editions of a
work into one doc. **Best use: resolve full EN author names from initial-only list entries**
(`title:"..."` → author_name) and get the EN works list + ISBNs to feed Google Books.
- **RU coverage is sparse but REAL** (romanized titles, e.g. «Chelovek i ego simvoly»); the
`search.json?isbn=` facet works for RU ISBNs too. Don't use a `language=rus` filter — it
under-reports. RU hunting still mainly RSL/libgen/flibusta/GBS SPb/Google Books.
### 12. Wikidata — identity anchors (life dates + OFFICIAL RU name)
- Tool: `tools/wd.py "Full Name"`
- `wbsearchentities` + `wbgetentities`: QID, P569/P570 (birth/death), EN/DE/RU labels, RU aliases.
The RU label is the single most reliable spelling for RSL/libgen queries.
- Only useful for Wikipedia-scale names (Jung, Neumann, Kandinsky, Jaffé, …); niche authors
are absent — fall back to OL title search or publisher pages.
### 13. archive.org — EN full texts: search, metadata, downloads ("English libgen"; user-confirmed 2026-07-16)
- Tool: `tools/ia.py search "creator:(Carl Gustav Jung)"` / `title "Aion"` / `meta <id>`
- `advancedsearch.php?q=...&output=json&fl[]=identifier,title,year,downloads,access-restricted-item,mediatype`
(Solr syntax; append `AND mediatype:texts` — creator: search is loose, includes images/misattributions).
- `archive.org/metadata/<id>` → file list (Text PDF / EPUB / DJVU / FB3 + sizes) — pick the file, then
direct download `/download/<item>/<file>` works (used for sec04 OCR verification; verified: Aion
`collectedworksof92cgju` = CW9/2 full text, open, 344k downloads).
- **Uses:** (1) EN full-text downloads for ✅ items (Jung CW volumes live here: `collectedworksof*`);
(2) EN identity/metadata; (3) content verification (OCR `_text.pdf` → pdftotext, e.g. the Neumann
4-essay match). `access-restricted-item: true` = controlled lending (borrow, no direct download).
- Item page: `archive.org/search?query=creator%3A%22C.+G.+Jung%22` (web UI of the same index).
- **`CarlJungCollectedWorks` item (OPEN, found 2026-07-18)** = Princeton CW set in one item: vols 1-18 (incl. 9/1, 9/2, 10=Kundalini 1996) + Jung Seminars 1 (Dream Analysis 1984), 520 (Zarathustra), 539 (Analytical Psychology 1925) + Children's Dreams 2012 + Zofingia Lectures + Synchronicity + Psychology of the Unconscious (Hinkle 1916). PDF + EPUB per volume, direct `/download/CarlJungCollectedWorks/<file>` works. Supersedes per-volume `collectedworksof*` items (many of those are borrow-only).
### 14. Google Books — book page via curl OR fetch_content (API is 429 without key)
- `books.google.com/books?vid=ISBN<13digits>`: **plain curl works** — 200 + `<title>"T - A -
Google Книги"` if indexed, 404 if not. Good ISBN validator incl. RU (АСТ/Эксмо/БукСМарт hit;
small-press 404). googleapis.com/books API stays 429 without a key.
- **Keep volume LOW**: Google may block agents on sustained querying (user warning 2026-07-15).
Budget: a few dozen requests per session, 2–5s random pauses, never in tight loops; treat as
an auxiliary cross-check, not a bulk source. If 429/403-redirect walls appear — stop and fall
back to OL/libgen/RSL.
- `fetch_content` on the same URL → richer metadata: full author/editor names, edition, series.
Resolved Ottmann=Klaus, Goldstein=Ralph, Brutsche=Paul, Elder=George, Moon=Beverly,
Killick=Katherine, Pennington/Staples, Rowland=Susan, Bolander=Karen, Acton=Mary via this or OL.
- **PagePlace preview PDFs** (2026-07-17): many Routledge/T&F books have a public preview at
`api.pageplace.de/preview/DT0400.<ISBN13>_A<n>/preview-<ISBN13>_A<n>.pdf` (~20–30 pp: title page +
TOC — perfect for identity/series verification when the full text is closed; find the URL via
SearXNG `"<EN title>" preview pdf`). Used for item 12 (Evers-Fahey, 9781317219583).
### 15. Shop directory (HSE bookshelf34 pattern) — where RU psychoanalytic books are sold
- HSE «Книжный шкаф» (hse.ru/ma/therapy/bookshelfNN) = a DIRECTORY of shops, not a book list:
NikBook (nikbook.ru — WBS shop, `?q=` search is fuzzy/works), Когито-Центр, **Скифия**
(skifiabook.ru — publisher/shop, JS catalog, has a full price list download), marketplaces.
- NikBook Jungian finds (2026-07-09): Тоцци «Активное воображение в теории, практике и обучении»,
Конгер «Юнг и Райх. Тело как Тень», Эдингер «Библия и психе».
- beta2alpha.ru — RU publisher (Tilda site; sells via Ozon seller beta-2-alpha); Jungian stock TBC.
### 16. Publisher sites — RU full texts and bibliographies
- hi-human.org (Living Human Heritage RU) — publishes book forewords/sections as pages (source
of the Neumann EN-works list with RU titles).
- castalia.ru / castaliasilvasacra.ru (Клуб Касталия) — author articles + shop.
- daimon-verlag.ch / chironpublications.com / spring-publications.com — EN Jungian presses
(spring-publications unreachable from this box; use Google Books ISBN instead).
### 17. alib.ru — RU/KZ book marketplace, 3.2M listings (used + new, incl. small press)
- Tool: `tools/alib.py "query"` (stem search! matches all inflections — unlike libgen).
- Endpoint: `https://www.alib.ru/find3.php4?tfind=<query>`; **query and pages are CP1251**
(utf8 → mojibake). Phrase mode: double quotes. Also: author-first, year range, ISBN digits,
price range. `>Купить<` links = listing count.
- Value: catches small-press / out-of-print editions shops don't carry (found: Furth 2nd ed.
2014, Turner handbook 2015, Neumann КДУ/Маниф/Питер editions). Seller titles can be wrong —
verify against RSL before trusting (e.g. «Челокес и миф» Kastaalia 2018 — not in castalia.ru
catalog → seller error).
- Also: alib.top (Ukraine), 33ob.ru (vinyl), amarka.ru (stamps) — sister sites.
### 18. Local CRW renderer (fastcrw.com, Firecrawl-compatible, localhost:3000)
- Tool: `tools/crw.py <url> [--html] [--wait N ms]` — POST /v1/scrape {"url","formats":["markdown"]}
- Use for JS-heavy sites curl can't read (soznanie.ast-academy.ru festival site rendered fine).
- Limits: Ozon (FAB challenge) and chitai-gorod SEARCH are API/anti-bot protected — CRW returns
nothing. chitai-gorod PRODUCT pages DO render via CRW (see source 9).
- Insales shops (castalia.ru) don't need it: product data is server-side JSON-LD (curl OK).
### 19. NLR / РНБ (National Library of Russia) via Primo (primo.nlr.ru) + CRW
- Tool: `tools/nlr.py "query"` — free-text search (all fields, words ANDed), page 1 (20 rows).
- Search is client-rendered → the tool renders the `search.do?fn=search&ct=search&
vl(freeText0)=QUERY&vid=07NLR_VU1&mode=Basic&initialSearch=true` URL via CRW and parses
title/author/year/holding (+ doc IDs when the linked layout renders).
- Full record (GOST description, incl. translators): CRW on
`display.do?tabs=detailsTab&ct=display&fn=search&doc=<07NLR_LMS#########>&displayMode=full&vid=07NLR_VU1`
(plain curl gives only the shell — details tab is XHR).
- Value: complements RSL — own St. Petersburg holdings + GOST records; author rows carry
life dates and NLR auth IDs. Count line parsing is flaky — trust the parsed row list.
- Verified 2026-07-09: all 35 section-04 surnames swept (see sections/04-pictures/SWEEP-R7.md).
### 20. RSL direct (aleph.rsl.ru, Ex Libris Aleph) — PARTIAL
- find-a flow works: GET /F/-?func=file&file_name=find-a → parse action URL (session token)
→ POST func=find-b&find_code=WAU&request=<name> → results 10/page with pagination.
- Full record: follow the result row's `full-set-set` link (session-dependent); gives
translator, series, ISBN, UDC (verified: Нойманн «Амур и Психея» МИФ 2024, ISBN
978-5-00214-519-5, пер. М. Виноградова).
- Quirk: find_code WTI (title) and SYS (ISBN) return EMPTY 0-byte responses; only WAU
(author) reliably works so far. libgen proxy (tools/rsl.py) still covers title/ISBN queries.
- search.rsl.ru — same catalog, Yii app, no JSON API found (2026-07-09).
### 21. ISBN validation pipeline — `tools/isbnval.py` (2026-07-15)
- Per ISBN: (1) libgen.vg biblioservice `type=isbn` — checksum + registered publisher group;
(2) Open Library `search.json?isbn=` — title/author/year; (3) Google Books `?vid=ISBN` curl
— title tag, 200/404; (4) NLR free-text for RU misses. Random pauses 2–5s, resumable.
- Verdicts: VALID (found w/ matching title) / STRUCT-ONLY (real ISBN, prefix registered,
not indexed — small-press RU) / BAD-STRUCT. Output TSV: `data/isbn-validation.tsv`.
- Found in sec04: 2 bad-check-digit typos (5-89613-003-6, 5-263-00368-2), 2 publisher-prefix
mismatches (Ленанд vs URSS; Академический проект vs Азбука-Аттикус), 2 co-publishing
prefixes (Петроглиф for ЦГИ 2020; Пальмира for Касталия 2025).
- **triumph.ru** (РГБ ISBN search, user-provided): `https://www.triumph.ru/html/serv/find-isbn.php?isbn=<13digits>`
(relative to the poisk-isbn.html page!) → returns links to РГБ/МГУ/РНБ/БЕН РАН/ГПНТБ catalogs
pre-filled with the ISBN. Itself no data; МГУ (nbmgu.ru) results are JS-rendered, БЕН РАН
(Koha) is scrapable but small collection. NLR absence ≠ ISBN invalid.
- **isbnsearch.org**: nice per-ISBN pages (author/publisher/year) but rate-walls after ~10
requests ("Please Verify to Continue", no recovery in 7 min) — not for batches.
### 22. MGU Scientific Library (nbmgu.ru) — server-rendered, curl-able (2026-07-15)
- Tool: `tools/mgu.py "query" [FIELD] [pages] [method]` or `--multi "AUT:X" "TIT:Y"`.
The "JS" advanced search is a plain GET: `/search/?adv=1&q1=<q>&f1=<F>&v1=<M>&cat=BOOK[&p=N]`.
- Fields: ANY/AUT/COA/TIT/KEY/RUB/YEA/PLA/PUB/SER/ISB/ISS/NBM; rows AND between q1..q3.
Method v: 0=Слова (stemming, default) · 1=Словосочетание · 2=Начинается с · 3=Дословно.
- "Всего: N" + GOST row (title/authors/notes) + "City : Publisher, Year" + shelf code + uid link.
20 rows/page, p=0-based; **out-of-range p silently falls back to page 0** (dedupe by uid!).
- **ISB field NOT populated** (0 hits, 3 formats × control ISBNs) — no ISBN search; use
AUT/TIT/SER/PUB. `storing.aspx?uid=` page = holdings/order only (no full record).
- Value: second RU catalog (complements RSL/РНБ), strong **SER** sweep (e.g. «Библиотека
аналитической психологии» → full series list in one query), GOST rows w/ translators.
- Anchor titles in result rows contain a raw `>` (title="Хранение<br/>Заказ") — regexes must
not use [^>]+ across the anchor attrs.
## Catalog caching convention
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source
URL + date in the header) so negative verification stays re-runnable without re-crawling.
Examples: `cogito-yungianskaya-series-2026-07-17.md` (30 pp. CRW crawl, 518 titles),
`litres-140409-2026-07-17.md`, `opp-misp-reading-list-2026-07-17.md`.
## Pipeline v3 (2026-07-17) — RU-title candidates & batch sweeps
- **`tools/rutitles.py`** — the "diff, not search" tool:
- `build` — rebuild `data/ru-titles.jsonl` from `data/catalogs/*` + preserve flibusta entries
- `addflib <authorId>` — append flibusta.is authorall titles to the DB
- `translate "EN Title"` — RU token bag via `data/ru-dict.tsv` (540+ EN→RU pairs, phrases first,
observed Kogito/Castalia/Peter conventions) + list of UNTRANSLATED words (finish manually)
- `match "RU tokens" [N]` / `both "EN Title"` — Jaccard over **pymorphy3-lemmatized** tokens
(case endings handled: «великой матери» → «Великая мать» 1.25; order-agnostic)
- DB (2026-07-17): 692 titles (Когито-518 + ЛитРес-24 + ОПП/МИСП-55 + flib a/5272, 57639, 118921,
122883, 193690, 193689)
- **Limits (honest):** generic RU titles (Wolff «Введение в основы…») invisible — publisher
sweep (Phase 2b) stays a separate phase. Translation ≠ identity: verify real hits.
- **pymorphy3** installed (user-site): lemmatization for RU token matching.
⚠ pymorphy2 is BROKEN on Python 3.12 (`inspect.getargspec` removed) — use pymorphy3 only.
- **`tools/oa.py`** — OpenAlex/Crossref (Phase 0, see source 10b).
- **flibusta .su dropped** from the search pipeline (mirror of .is, reduced functionality).
- **`tools/sweep.py`** — batch driver for Phases 1–2 (per author/item: rsl + lg + alib + flib +
cogito + SearXNG-OZON-snippet parse + rutitles diff; optional `--nlr` for НРБ via CRW, slow).
Query log: `data/queries.log` (JSONL). Sweep output: `data/sweeps/secNN/<nn>-<slug>.md`.
## Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)
- bookmate.com (403 bot protection), gtmarket.ru (JS app, no HTML content)
- other libgen mirrors (libgen.lol → error page; .vg works with browser UA)
## Identity check protocol (MANDATORY before trusting a hit)
1. **Life dates** in RSL author field (e.g. `Нойманн, Эрих [1905-1960]`).
2. **EN↔RU cross-mapping:** the found person's *other* RU books must map to known EN titles
(e.g. found «Страх феминного» + «Любовь и Душа» → EN *Feminine* + *Love and its Opposites*
→ same Erich Neumann).
3. **Co-publisher signature:** known translation pipelines (Vakler/Рефл-бук 1990s,
Касталия, Прайм-Еврознак, АСТ/ЭКСМО, Питер «Мастера психологии»).
4. **Content spot-check** of the actual target title (open edition page / TOC) when equating
an RU title with an EN title that sounds different
(e.g. *The Archetypal World of Henry Moore* ⇄ «Искусство и творческое бессознательное» — verify it is about Moore).
5. **Homonym rejection:** different first name / field / era → reject or mark ambiguous
(trap found: Верена Каст ≠ Вернер Каст; «Каст» search returns Verena Kast books).
## Workflow (per section) — v2 (2026-07-09)
**Phase 0 — Full names + identity** (`authors/`):
1. Resolve every initial-only author to a full EN name: `tools/ol.py 'title:"EN Title"'`
(author_name) → else `tools/wd.py` (QID + RU label) → else Google Books ISBN page →
else publisher/review page. Write/extend the **dossier** `authors/<slug>.md` (EN full
name, dates, RU name to query, name source); then add a one-line entry in the thin
index `sections/<n>/AUTHORS-EN.md` linking to it. If the dossier already exists from
another section, just append the new list item + link it.
2. For big names: `wd.py` gives life dates + the OFFICIAL RU label (best RSL spelling).
3. Build the transliteration matrix per author (2–3 variants: Нойманн/Нойман, Кох/Коч,
Калфф/Кальфф, Ферс/Фёрт) — run every query with ALL variants.
**Phase 1 — Author sweep (EN→RU):** per author: `rsl.py "<First> <Last>"` → `lg.py "<Last>"`
→ `cogito q=<Last>` → `flibusta ?ask=<Last>` → livelib/chitai-gorod author page. Apply identity
check. Fill author dossier.
**Phase 2 — Title-gap sweep** for still-unmatched books: `rsl.py "<RU title words>"` (check
author field!), `lg.py` 2–3 whole-word title combos, cogito, flibusta by RU title, GBS SPb via
SearXNG `site:gbs.spb.ru`.
**Phase 2b — Publisher sweep (RU→EN, reverse direction):** per Jungian RU publisher/imprint
(Касталия, hi-human/Living Human Heritage, ЦГИ/Т8, Когито, Фантом Пресс, Азбука, Серебряные
нити, Питер, Ваклер, Добросвет, Деметра, НикБук, Скифия, beta2alpha, Эксмо/АСТ психология):
enumerate the full Jungian catalog (site search or RSL **series** queries: «Юнгианская
психология», «Библиотека аналитической психологии», «Классики зарубежной психологии») → map
every title to an EN original → backfill both missing items and future sections. This catches
books author queries can never hit (generic titles like «Сборник статей», reworked titles).
**Phase 2c — Community reading lists (oracles):** ISAP Zurich, HSE (hse.ru/ma/therapy),
МИСП, Jung Institute Moscow, education-psy.ru sandplay list, carljung.ru, koob.ru — these
explicitly state which books have RU editions (or lack them). Scrape each, diff vs our ❌ list.
**Phase 3 — Content verification (for suspicious candidates):** when an RU title's mapping to
an EN title is not obvious, verify 1:1 against the EN original: archive.org OCR text (TOC /
editorial note) vs RU essay/chapter list (koob/livelib/preview). Example: Neumann «Сборник
статей» = Art and the Creative Unconscious: Four Essays (4/4 essay match).
**Phase 4 — Edition collection:** ALL editions from ALL sources (title variants!), full metadata
(publisher/year/pages/translator/ISBN) — prefer RSL/GBS SPb records (GOST) over shop data.
**Phase 5 — Book notes + commits** (one commit per completed book or meaningful author gain).
When a section is complete, its `SUMMARY.md` must contain a consolidated
**"Publication status at a glance"** table (one row per reading-list item) with these columns:
`# | EN book | EN ISBN | RU book | RU ISBN | DL EN | DL RU`. This is the single
quick-reference view — it records not just ISBNs but, for each item, whether the EN original
and the RU translation are actually **downloaded** (file count + format, e.g. "5 pdf", "2 fb2",
"1 djvu", "1 doc", or "sec04: N pdf" for cross-section items, or "— (restricted)" / "— (absent)" when
not downloadable). Legend for ISBN columns: **✅** = ISBN found; **⚠** = published but no ISBN
(not stated / new ed. / OZON-prefix only / truncated); **❌** = no such edition (verified absent);
**—** = the book itself has no such edition (essay / Standard Ed. vol / pre-ISBN). End the table
with a one-line totals row (DL EN / DL RU file counts). Cross-check the DL columns against the
section's `downloads/<nn>-*/MANIFEST.md` so the two never disagree.
Negatives gate (before ❌): author sweep (RSL + libgen + cogito + flibusta) + one title-level
SearXNG + livelib author page + one shop/retail check. Log every RU query in MISSING.md.
Special cases:
- **Jung's CW volumes:** translated and widely available — don't deep-search; note the RU
volume for the item (CW15 = «Дух в человеке, искусстве и литературе», Харвест 2003;
Red Book = «Красная книга (Liber Novus)», Касталия 2025, trans. О. Комков).
- **Essays/chapters** (e.g. Jaffé in *Man and his Symbols*): mark 🔶 if only the parent
book/anthology is available in RU.
## Politeness
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.
## Download phase conventions (user rule, 2026-07-18)
- **EN originals: libgen.vg FIRST** (the EN collection is bigger than the RU one — most trade
academic books: Routledge/Karnac/SUNY/Princeton/etc.), **archive.org second** (CW/Philemon
open scans, content-verification OCR, items libgen lacks). Flibusta for EN: tried
2026-07-18 — OPDS search is effectively RU-only (EN titles 0, author search returns only
Cyrillic pages); a one-shot exact-title probe is cheap, don't plan on it.
- **RU translations: libgen.vg + flibusta.is** (flib = fb2/epub when libgen has only scan-pdf or nothing).
- Per language, format priority: ALL pdfs if any pdf exists; else all fb2; else all epub; else all doc.
- File naming `NN-<slug>-<lang>-<edition>.<ext>`; `downloads/0N-<name>/` gitignored, only `MANIFEST.md` committed.