jung/AGENTS.md
Dmitry Kokorin bfba5fa048 sweep 01-06 wrap-up: READING-LIST regen (583 files), generator xref-tally fix, INDEX fixes
- make_rootlist.py: xref items no longer counted into section status tallies
  (they carry the canonical card's marker in parens, which was inflating counts:
  sec11 6/1/0, sec10 5/0/27 now correct)
- sec10 INDEX: #23 Jamison, #33 Saks ❌->✅ (stale after user-oracle rescues)
- sec03 MISSING: #65 stale entry counted out (already corrected 2026-09-24)
- AGENTS.md: section file counts + sec03 40/0/28 + total 583 files
- READING-LIST.md regenerated: 583 files, 768 links
2026-09-29 10:10:20 +03:00

834 lines
70 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Jung Reading List → Russian Editions: Research Notes
## Goal
For every book in the ISAP Zurich reading list (https://isapzurich.com/en/library/reading-list),
find the **Russian edition(s)**: correct Russian title, publisher, year, translator, ISBN,
libgen download page (if exists), and РГБ (RSL) bibliographic record (if exists).
We work **section by section**. Status (детали — в SUMMARY.md/MANIFEST.md секции + git log):
| sec | name | items | RU ✅/🔶/❌ | files (EN+RU) | state |
|-----|------|-------|-----------|----------|-------|
| 01 | fundamentals | 55 | 38/2/15 | 114 | FINAL 2026-09-25 (+1 EN 2026-09-28) |
| 02 | dreams | 57 | 23/5/29 | 38 | FINAL 2026-09-25 (+1 EN 2026-09-28) |
| 03 | myths & fairy tales | 68 | 40/0/28 | 144 | FINAL 2026-09-25 (+14 EN 2026-09-28) |
| 04 | pictures | 39 | 12/1/26 | 29 | FINAL 2026-09-25 (+1 EN 2026-09-28) |
| 05 | ethnology | 29 | 17/0/12 | 36 | FINAL 2026-09-25 (−2 → sec03, 2026-09-28) |
| 06 | religion | 28 | 19/1/8 | 38 | FINAL 2026-09-25 |
| 07 | complexes | 27 | 4/0/23 | 20 (EN only) | FINAL 2026-09-26 |
| 08 | developmental | 48 | 20/1/27 | 44 (333 MB) | FINAL 2026-09-26 |
| 09 | comparison of psychodynamic concepts | 38 (20 NEW + 18 xref) | 11/1/8 | 43 (366 MB) | FINAL 2026-09-27 |
| 10 | psychopathology & psychiatry | 37 (32 NEW + 5 xref) | 5/0/27 | 30 (268 MB) | FINAL 2026-09-27 |
| 11 | individuation process | 22 (7 NEW + 15 xref) | 6/1/0 | 10 (128 MB) | FINAL 2026-09-27 |
| 12 | practical case | 35 (23 NEW + 12 xref) | 14/1/20 | 35 (184 MB) | **FINAL 2026-09-28** |
**ALL 12 SECTIONS COMPLETE (2026-09-28).** Total: 483 positions, 583 files (~2.5 GB; +16 EN sweep 01-06, 2026-09-28/29). Root view: `READING-LIST.md` (generated by `tools/make_rootlist.py`; regenerable, do not hand-edit).
**PROJECT COMPLETE (2026-09-28).** All 12 sections final. Follow-ups: user-oracle re-probes on ❌ lists (per-section MISSING.md), AGENTS.md variant A (extract sources).
## Repository layout
```
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
data/MASTER-LIST.md # CANONICAL: all 12 sections, cross-section dedup (tools/xref.py --write)
READING-LIST.md # ROOT VIEW: весь список как на сайте ISAP + ссылки на файлы каждой книги
# (генерируется tools/make_rootlist.py из INDEX + MANIFEST + карточек; не редактировать руками)
data/isap-raw/ # RAW-списки ISAP 08–12 (scope pending); 07 в sections/07-complexes/RAW.md
tools/ # search helpers (lg.py, rsl.py, sx.sh, xref.py, audit_anchors.py, …)
authors/ # CANONICAL: one dossier per author (shared across all sections)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
# + AUTHORS-EN.md = THIN INDEX (item → author → link to authors/<slug>.md)
```
## Conventions
- **One file per book**, named `NN-<slug>.md` (NN = order in the reading list).
- **Status marker** at top of each book file:
- `⬜ not searched yet`
- `🔎 searching in progress`
- `✅ RU edition(s) found`
- `🔶 partial` (e.g. only an essay/chapter available in RU)
- `❌ no RU edition found (verified)`
- **Book card CANON v3 (agreed 2026-09-21; mandatory for new cards, sec05+ from day one):**
```markdown
# <listed title> (<listed edition: publisher, year>)
**Author(s):** <Last, First> [dates]
**Shelf mark:** <ISAP or —>
**Section:** <NN name / subgroup>
**Original:** <язык оригинала, изд., год — только если ≠ языка списка (ES/DE/…)>
**Status:** <✅ | 🔶 | ❌> — <одна фраза: «RU 4 изд.» / «essay only» / «no RU (verified <date>)»>
## Editions — EN ← язык списка; несколько строк = несколько изданий
| Title / edition | Publisher | Year | Pages | ISBN | Notes |
|-----------------|-----------|------|-------|------|-------|
## Editions — RU ← тот же формат + колонка Translator
| RU title | Publisher | Year | Pages | Translator | ISBN |
|----------|-----------|------|-------|------------|------|
*(нет: одна строка `— нет (verified <date>, gate: MISSING.md §NN)`)*
## Editions — ES ← только если есть изд./оригинал на 3-м языке
## Downloads ← только если есть файлы; симметрично по языкам
| lang | file | size | source |
|------|------|------|--------|
| EN | NN-…-en-<edition>.pdf | 64 MB | lg f/… | ← source: lg f/ | flib b/ | ia <item>
| RU | NN-…-ru-<edition>.fb2 | 4.4 MB | flib b/… |
## Catalog records
РГБ: rsl.py "<query>": [id] <GOST-строка>; [id2] …
РНБ: ✓ 3 (только если sweep трогал; добавления — с данными)
МГУ: ✓ 1 / GBS: ✗ 0 / flib: a/<id> (N книг)
EN refs: OL (N ed., ISBNs ✓ isbnval) · Google Books 200 · ia: <item>
*(❌-карточка — ОДИН compact gate: `РГБ 0 · НРБ 0 · МГУ 0 · flib 0 (verified <date>, gate: MISSING.md §NN)`)*
## Notes
- <идентичность, identity-check, cross-refs на досье/секции; датированные append: «2026-09-20: …»>
```
Rules: (a) Editions-секции симметричны по языкам, порядок: язык списка → RU → прочие;
(b) ISBN только в таблицах (из шапки убрано поле EN ISBN); (c) пустые секции не пишутся;
(d) ЗАПРЕЩЕНЫ плейсхолдеры: `// not searched yet`, `(empty)`, дубли `## Notes`;
(e) `## Libgen` более нет (был асимметричным); файлы ≠ издания (3 файла одного изд. = 3 строки);
(f) EN-каталогами: полные записи не ведутся (WorldCat мёртв) — строка `EN refs` по данным
isbnval/OL/ia; (g) ✅-карточка: РГБ обязателен + ≥1 второй каталог если sweep его трогал;
(h) старое поле Status-формулировок: единая «маркер + одна фраза».
Все карточки мигрированы на v3 (2026-09-21, one-shot `tools/migrate_cards_v3.py`);
linter `tools/check-md.py` = 0 issues. New cards start on v3.
- **Section files unified 2026-09-21** (CANONs ниже; one-shot скрипты `tools/migrate_*.py`,
`make_authors_table.py`, `restructure_summary.py` — done, не для повторного запуска).
Living authors keep open-ended dates `[1951-]`.
- **Section file CANONS v1 (agreed 2026-09-21):**
**SUMMARY.md:** H1 `# Section NN — <Name>: summary (<date>)` → `## Final counts` (одна строка:
статусы + файлы) → `## Publication status at a glance` (ОБЯЗАТЕЛЬНЫЙ, колонки `# | EN book | EN ISBN
| RU book | RU ISBN | DL EN | DL RU` + totals-строка + legend ✅/⚠/❌) → `## Authors — EN → RU`
(col `item(s) | EN author | RU name | verified via`) → `## Notes` (датированные append).
Запрещено: «Table N»-нумерация, «Table 2 — Books» (дубль карточек), ISBN-validation-таблицы
(данные в data/isbn-validation.tsv, строка-ссылка).
**INDEX.md:** одна таблица `| # | Group | EN title | Author(s) | Status | File |` (Group = A.1/B.1,
нет подгрупп → `—`; Status = только маркер; File = имя карточки).
**MANIFEST.md** (`downloads/NN-<name>/`): H1 `# Downloads — Section NN <Name> — FINAL <date>` →
`## Totals` (N files, X GB, RU/EN split) → `## Files (per item)`: на каждый item `### NN — <title>
— <status>` + таблица `| file | size | source | verify |` (source: lg f/… | flib b/… | ia <item>).
**Cross-ref rule (2026-09-26, user):** item, чья книга впервые появилась в предыдущей секции,
ОБЯЗАТЕЛЬНО получает свой per-item блок в основном списке (заголовок «— files in sec0N» +
одна cross-ref-строка `| — (cross-ref sec0X #NN: <файлы>) | … |`); блок `## Cross-references`
= только резюме-указатель, не замена per-item блокам.
→ `## Not downloadable (verified <date>)` таблица `| item | reason |` → `## Cross-section` → `## Notes`.
**MISSING.md:** H1 → `## ❌ (N)` → на каждый `### NN — <EN title> (<author>)` + 1–3 строки evidence
(каталоги 0, verified date) — `### NN` = ЯКОРЬ для gate-строк карточек (`gate: MISSING.md §NN`)
→ `## 🔶 (N)` → `## Re-probes` (датированные append).
**Dossiers** (`authors/*.md`): H1 `# <Lastname, Firstname> [<dates>]` → bold-поля `**Wiki anchor:**` / `**Dates:** /
**Field / identity:** / **RU name (canonical):** / **RU name (candidates):**` → `## List items
(ISAP Zurich)` таблица `| sec | # | title | status |` → `## Verified RU works` таблица
`| RU title | EN original | publisher, year | ISBN |` → `## Homonym warnings` (ЕДИНОЕ название,
legacy-варианты: «Homonym warning (date)», «Homonyms / notes», «Омонимы») → `## Notes` (датированные
append; legacy «Round N» логи сюда). Запрещено: `**EN identity:**` отдельно, `**Dossier status:**`.
**Wiki anchor (2026-09-24, MANDATORY):** `**Wiki anchor:** QID (URL)` — факты identity (даты,
родство, поле) БЕЗ якоря = unverified и не пишутся как утверждение. Якорь = Wikidata QID
(wbgetentities по QID — точечно, без омонимов) + en/ru-wiki URL. Инструмент:
`tools/verify_dossiers.py` (resumable, TSV `data/dossier-check.tsv`, 3s/polite + 429-backoff;
статусы: MATCH = QID сверен, CANDIDATE = найден поиском — РЕВЬЮ перед записью якоря,
NO-QID = legitimately absent). Предыстория: 2 досье с галлюцинациями (Emma Jung: «1877-1965» +
«eldest daughter» вместо 1882-1955 + жена; Verena Kast: «1951» + «daughter of Hans Kast» вместо
1943 + отец Walter Kast) — lesson: identity-факты только с источником.
- **TODO (later, after 4 sections in order):** `tools/liart.py` — **HALF DONE 2026-09-21:**
OPAC-Global protocol reverse-engineered (direct.exe/FindView, guest session), SearXNG
record-URL fallback works; BLOCKED on search labels + iddb (server InfoDB, admin-only for
GUEST) — needs a one-time browser read (see data/liart-params.json).
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) are the **CANONICAL** home for
identity, shared across ALL sections (same author can appear in 01, 04, …). One dossier per
PERSON, slug `<lastname>-<firstname>` (lowercase, hyphens). Hold: dates, RU name candidates
(all plausible transliterations, German names often have several: Нойманн/Нойман/Нейманн),
the list items that cite this author (sec # + status), verified RU works with EN mapping,
homonym warnings.
- **`sections/0N/AUTHORS-EN.md` is a THIN INDEX only** — table `item(s) → author → link to
`../../authors/<slug>.md``. It must NOT duplicate identity data (no RU-name analysis, no RU
works table). If an author appears in two sections, the dossier is written once and both
section indexes point to it. Cross-section homonym traps (e.g. the two different "Alice
Miller"s) get their own dossier + a warning in both.
- **Commit per completed book** (or when meaningful author info is gained), message like:
`sec04: <book> — <finding>`.
- Update `CONTEXT.md` for anything useful that fits no category.
## Sources and their status (verified 2026)
### 1. РГБ (Russian State Library) via libgen.vg biblioservice — PRIMARY for existence + editions
- Tool: `tools/rsl.py "Full Name"` (JSON, polite, 3s between pages)
- Raw URL: `https://libgen.vg/biblioservice.php?value=<query>&type=rsl&format=json`
- Rich records: title parts, **author with life dates** `[1905-1960]` (great identity anchor),
translator, publisher, city, year, pages, ISBN, series, UDC tags, even TOC field.
- **Limitations:** no pagination (fixed first page) → always query by **full first+last name**
for people, or 2–3 distinctive words. Loose matching (searches author AND title fields).
- Other `type=` values on this mirror are **dead**: worldcat, googlebooks, isbndb, udc
(all return "Nothing found" even for control queries). `type=isbn` decodes ISBN structure:
checksum error + **registered publisher group name** (e.g. 5-519 → «Клуб Касталия»,
5-98712 → «Петроглиф») — cheap, unlimited; used by `tools/isbnval.py` as the structure layer.
### 2. libgen.vg full-text search — PRIMARY for readable editions
- Tool: `tools/lg.py "query"` (whole words only, res=100, prints pagination hint)
- Raw URL: `https://libgen.vg/index.php?req=<query>&res=100`
- **Requires a browser User-Agent** (otherwise nginx default page / robot block).
- **Search semantics (verified):** whole-word, case-insensitive. Stems do NOT work reliably
(`происхожден` → 0 hits; full `Происхождение` → many). Multi-word = AND.
→ Use **full inflected words** (nominative, as stored in titles) and 2–3 distinctive words.
- Author search works (`req=Нойманн`) but pulls homonyms — filter by first name in results.
- Pagination: `&page=N` works.
- **Search objects (2026-09-25):** `objects[]` param = что искать (f=files, e=editions, s=series…).
Без него = editions-only — файлы, не привязанные к нужному изданию, НЕ находятся
(Solomon: editions-поиск = только mislinked 2021-издание, files-поиск = настоящий файл 2018-го).
lg.py теперь шлёт `objects[]=f,e,s,a,p,w` + `topics[]=l,c,f,a,m,r,s` и печатает `[md5=…]` в строках
→ `tools/lgdl.py bymd5 <md5>` = f_id + метаданные. Это нашло все 6 rescues sec06.
- Edition page: `https://libgen.vg/edition.php?id=<id>`.
### 3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
- URL: `https://cogito-shop.com/search/?q=<query>` (param is **q**, not query)
- Matches author AND title (fuzzy — «исцеляющее сновидение» missed, «Майер» hit). Good for
"is it sellable now in RU".
- **Product pages work with plain curl** (2026-09-19): «Характеристики» block in HTML —
Автор / Издательство / Формат / Вес / Тип обложки / Кол-во стр / Год / ISBN / Код. Search URL
= `/search/?q=<author>`, product link = `/catalog/<cat>/<slug>/`.
### 4. SearXNG (local) — cross-verification
- Tool: `tools/sx.sh "query"` (JSON API at http://localhost:8888; the web_search tool blocks localhost)
- Use to verify RU↔EN title pairing in a third party's words, resolve transliterations.
- Keep queries ≤ 3–4 words; «…» quotes only on distinctive RU titles. Unstable relevance —
drop junk queries, don't retry blindly.
- **Engine status (2026-09-19):** container `searxng` (docker, config
`~/.searxng/settings.yml` → mounted /etc/searxng/settings.yml). Currently **google only**:
bing is BLOCKED (responds but ignores query — random promo junk), brave dead (403),
duckduckgo CAPTCHA, wikipedia/wikidata dead/timeout — all disabled in config.
Diagnosis: `docker logs searxng | grep -oE '(ERROR|WARNING):searx.(engines|network).[a-z_]+' | sort | uniq -c`
+ per-engine test `curl --get :8888/search --data-urlencode 'q=...' --data-urlencode format=json --data-urlencode 'engines=google'`.
Fix = mark engine `disabled: true` in settings.yml + `docker restart searxng`.
- **Always pass explicit `language=ru` / `language=en`** (user-confirmed fix; per-request, config stays multilingual).
### 5. ru.wikipedia / en.wikipedia — author dossiers
- ru-wiki articles often have «…на русском» sections listing actual RU titles.
- en-wiki for the EN bibliography (identity anchor: 3–5 flagship titles).
### 6. livelib.ru author pages — all RU works of a known author
- e.g. `livelib.ru/author/<id>-<slug>` lists every RU edition (server-rendered, scrapable).
- Author ID: find via SearXNG `site:livelib.ru <author>`. The /search page is JS-only (nope).
- **Publisher pages** `/publisher/<id>-<slug>` = the publisher's full book list (server-rendered,
book links `/book/<id>-<slug>` carry title + author slugs) — cheap reverse publisher sweep.
Example: /publisher/2076-medkov-s-b (Медков С. Б. — small esoteric press w/ Jung line).
### 7. cogito-shop person pages — shop stock per author
- `cogito-shop.com/person/<first>_<last>/` (transliterated, first-name first,
e.g. `/person/mariya_luiza_fon_frants/`, `/person/yaffe_aniela/`).
- Slug discoverable from the shop's search page (results contain person links) or guessable.
- Lists every edition the shop carries (title variants!). Product pages lack full
metadata in HTML (JS-rendered) — use RSL/livelib for publisher/year/ISBN.
### 8. GNB SPb (State Public Library of St. Petersburg) catalog — SECONDARY bibliographic source
- URL: `https://www.gbs.spb.ru/ru/search/detail/?id=<hash>` (rich GOST records: title,
responsible parties incl. translators, publisher w/ full legal names, year, pages, ISBN,
notes like «Др. кн. авт.»)
- **Complements RSL:** RSL (RGB) misses some small-press editions (e.g. Pattis-Zhoya
«Аборты…» Т8/ЦГИ 2017 is in GBS but NOT in RSL). Query via SearXNG `site:gbs.spb.ru <author>`.
### 9. chitai-gorod.ru / ozon / labirint — trade retail pages
- chitai-gorod author pages list RU works (some noise). **Product pages: CRW works** (2026-09-19) —
«Характеристики» block: Год издания / Кол-во стр / Переводчик / Издательство / ISBN. Plain curl
gives nothing (JS). Find product IDs via SearXNG `site:chitai-gorod.ru "<RU title>"`.
- labirint.ru book pages work with plain curl — «Характеристики» block after <h1> (publisher,
year, pages, translator, ISBN). Labirint SEARCH is anti-bot (text= ignored).
- Ozon product pages: blocked from curl (redirect loop) AND via CRW (FAB challenge 2026-09-19);
fetch_content also fails. Use SearXNG OZON snippets instead.
- vse-svobodny.com — works with `Cookie: beget=begetok` (anti-bot challenge, 2026-09-19);
product pages carry publisher/year/pages.
#### 10. flibusta — big e-book library. .su = server-rendered search; **.is = OPDS + direct downloads (user fixed access, 2026-07-16)**
- **.is OPDS (preferred):** Tool `tools/flib.py`:
- `flib.py authors "query"` → `/opds/search?searchType=authors&searchTerm=…` (id, name, book count)
- `flib.py books "query"` → `searchType=books` (id, title, author, year, format, translator from annotation)
- `flib.py author <id>` → `/opds/author/<id>/alphabet` — book list (**20 per page!**)
- `flib.py authorall <id>` — follows `rel="next"` pages (`/alphabet/1`, `/2`, …) to the end — **always use authorall** (Jung a/5272 = 128 books; the bare feed shows only the first 20)
- Author HTML: `/a/<id>` (e.g. a/57639 = фон Франц, a/118921 = Нойманн); book page `/b/<id>`
- **Direct downloads:** `/b/<id>/fb2|epub|mobi|html|txt` (pdf/doc as «/b/<id>/download» links on the pages)
- **Quirks (2026-07-17):** `fb2+zip` format actually returns a ZIP (unpack: 1 fb2 + cover inside); `/b/<id>/pdf` sometimes 302-redirects to the book HTML page (a ~20 KB stub) when the file isn't served directly — check downloaded «pdf»s with `file`, fall back to libgen (same editions usually there) or the fb2 twin.
- **Converter quirk (2026-07-18):** `/b/<id>/epub` (or any fmt endpoint) on a book whose NATIVE format is pdf/djvu/doc may return the **native file, not an epub** — always magic-byte-check downloads (`file`); some «20 KB HTML stubs» are in fact real native files with the wrong extension.
- .su = **mirror of .is with reduced functionality** (user, 2026-07-17) — use .is OPDS only.
.su `booksearch/?ask=` exists but adds nothing; API endpoints blocked (api/search.php → 403).
- Value: (1) RU full-text downloads for the download phase; (2) author-book-list = cheap author sweep;
(3) negative verification (all same-name authors listed). Verified 2026-07-09 (.su): no «Выявляющий
образ», no «Сэндплей» titles, no «Тест Дерево», no «Искусство и творческое бессознательное» (as titled).
Flibusta search is fuzzy (whole-phrase not required).
### 10b. OpenAlex + Crossref (`tools/oa.py`) — Phase 0 identity/ISBN (free, no key, better than OL for modern academic books)
- `oa.py "Exact EN Title"` (OpenAlex works: full author names, year, DOI, publisher),
`oa.py --crossref "Title words"` (Crossref: **ISBN list** + publisher + year),
`oa.py --author "Last, First"` (OpenAlex author entity), `oa.py both "Title"`.
- Verified 2026-07-17: Evers-Fahey full name "Karen Evers-Fahey" + Routledge 2016;
"Coming into Mind" = **Margaret** Wilkinson (not Michael!) 9781317710578;
"The Symbolic Quest" Whitmont 9780691213187 (Princeton).
- Use in Phase 0 INSTEAD of / alongside OL; OL still primary for pre-1990 titles.
### 10c. DOI pipeline for article items (`tools/doi.py`) — Crossref → DOI → libgen (2026-09-26, user idea)
- Reading-list items that are JOURNAL ARTICLES: `doi.py "Last" "Title words" --year YYYY`
= Crossref (query.bibliographic + query.author, year ±2) → top-DOI hits ranked by title-token
overlap → for each: **libgen `req=<DOI>`** (libgen indexes papers BY DOI — verified: Krieger
10.1111/1468-5922.12544 → exactly ed 84462696; Meier I 10.1111/1468-5922.12545 → 84462697;
Bovensiepen 10.1111/j.0021-8774.2006.00602.x → 15358298; controls 3/3).
- Then `lgdl.py dl <f_id> ...` + **content-verify: first page must be the PAPER (author +
journal + pages matching RAW + abstract), NOT a review of it** (user rule 2026-09-26).
- Limits: pre-1997 works have no DOIs (1975 Hill paper — fall back to title sweeps);
Crossref top hits can be homonym noise (score column disambiguates; pick the [0.8x] Jungian one).
- Recovered sec07 EN via this route: #05, #17 (ed from user), #20. Also found a JAP 2008
Bovensiepen book-REVIEW (10.1111/j.1468-5922.2008.00737_3.x) — the review-check matters.
### 10d. flibusta.**su** — the OTHER flibusta (different content, 2026-09-26, user-found)
- flibusta.su has books NOT on flibusta.is (unique subset). Verified: CW3 RU
«Психогенез душевных болезней» (АСТ 2025, 978-5-17-139315-1) = b/390112 on .su, ABSENT on .is
(a/5272 has 129 books, no CW3). So **.is-only author sweeps can miss RU editions** (sec07 CW3 gap).
- Book/author pages curl-able (server-rendered titles). Downloads = JS POST `/lang/?b=<id>`
→ `"b":"0"` = partner-locked (litres exclusive, no file) or a file domain. No static `/b/<id>/<fmt>`
on .su (that's .is). For .su-only books, source the file via libgen/RSL by ISBN, not .su.
- Jung .su author page: 31 books (a/877 is the WRONG different-person «Бартольд»; the Jung .su
author id is unknown — find via the book page's author link).
### 11. Open Library — the "English RSL" (author works lists, full names, ISBNs)
- Tool: `tools/ol.py "author:Lastname, First"` or `tools/ol.py 'title:"Exact EN Title"'`
- `openlibrary.org/search.json?q=...&fields=title,author_name,publish_year,language,publisher,isbn`
- No key, no rate limits observed (1s between calls is polite). Aggregates all editions of a
work into one doc. **Best use: resolve full EN author names from initial-only list entries**
(`title:"..."` → author_name) and get the EN works list + ISBNs to feed Google Books.
- **RU coverage is sparse but REAL** (romanized titles, e.g. «Chelovek i ego simvoly»); the
`search.json?isbn=` facet works for RU ISBNs too. Don't use a `language=rus` filter — it
under-reports. RU hunting still mainly RSL/libgen/flibusta/GBS SPb/Google Books.
### 12. Wikidata — identity anchors (life dates + OFFICIAL RU name)
- Tool: `tools/wd.py "Full Name"`
- `wbsearchentities` + `wbgetentities`: QID, P569/P570 (birth/death), EN/DE/RU labels, RU aliases.
The RU label is the single most reliable spelling for RSL/libgen queries.
- **`tools/rulabel.py` (2026-09-24):** QID → ru label + ru aliases (canonical Cyrillic spelling).
Authority-DB route (VIAF/ISNI/GND) verified DEAD from this box: VIAF 403 + JS-app, ISNI 403,
GND API gone (services.dnb.de 404s), BnF SRU 403 — Wikidata ru labels are the working channel
(user idea, 2026-09-24). Real catches: Zimmer = «Циммер» not «Зиммер», Corbin = «Корбен» not «Корбин».
Niche authors may have NO ru label (MacCulloch) — fall back to RSL author field.
- Only useful for Wikipedia-scale names (Jung, Neumann, Kandinsky, Jaffé, …); niche authors
are absent — fall back to OL title search or publisher pages.
### 13. archive.org — EN full texts: search, metadata, downloads ("English libgen"; user-confirmed 2026-07-16)
- Tool: `tools/ia.py search "creator:(Carl Gustav Jung)"` / `title "Aion"` / `meta <id>`
- `advancedsearch.php?q=...&output=json&fl[]=identifier,title,year,downloads,access-restricted-item,mediatype`
(Solr syntax; append `AND mediatype:texts` — creator: search is loose, includes images/misattributions).
- `archive.org/metadata/<id>` → file list (Text PDF / EPUB / DJVU / FB3 + sizes) — pick the file, then
direct download `/download/<item>/<file>` works (used for sec04 OCR verification; verified: Aion
`collectedworksof92cgju` = CW9/2 full text, open, 344k downloads).
- **Uses:** (1) EN full-text downloads for ✅ items (Jung CW volumes live here: `collectedworksof*`);
(2) EN identity/metadata; (3) content verification (OCR `_text.pdf` → pdftotext, e.g. the Neumann
4-essay match). `access-restricted-item: true` = controlled lending (borrow, no direct download).
- Item page: `archive.org/search?query=creator%3A%22C.+G.+Jung%22` (web UI of the same index).
- **`CarlJungCollectedWorks` item (OPEN, found 2026-07-18)** = Princeton CW set in one item: vols 1-18 (incl. 9/1, 9/2, 10=Kundalini 1996) + Jung Seminars 1 (Dream Analysis 1984), 520 (Zarathustra), 539 (Analytical Psychology 1925) + Children's Dreams 2012 + Zofingia Lectures + Synchronicity + Psychology of the Unconscious (Hinkle 1916). PDF + EPUB per volume, direct `/download/CarlJungCollectedWorks/<file>` works. Supersedes per-volume `collectedworksof*` items (many of those are borrow-only).
### 14. Google Books — book page via curl OR fetch_content (API is 429 without key)
- `books.google.com/books?vid=ISBN<13digits>`: **plain curl works** — 200 + `<title>"T - A -
Google Книги"` if indexed, 404 if not. Good ISBN validator incl. RU (АСТ/Эксмо/БукСМарт hit;
small-press 404). googleapis.com/books API stays 429 without a key.
- **Keep volume LOW**: Google may block agents on sustained querying (user warning 2026-07-15).
Budget: a few dozen requests per session, 2–5s random pauses, never in tight loops; treat as
an auxiliary cross-check, not a bulk source. If 429/403-redirect walls appear — stop and fall
back to OL/libgen/RSL.
- `fetch_content` on the same URL → richer metadata: full author/editor names, edition, series.
Resolved Ottmann=Klaus, Goldstein=Ralph, Brutsche=Paul, Elder=George, Moon=Beverly,
Killick=Katherine, Pennington/Staples, Rowland=Susan, Bolander=Karen, Acton=Mary via this or OL.
- **PagePlace preview PDFs** (2026-07-17): many Routledge/T&F books have a public preview at
`api.pageplace.de/preview/DT0400.<ISBN13>_A<n>/preview-<ISBN13>_A<n>.pdf` (~20–30 pp: title page +
TOC — perfect for identity/series verification when the full text is closed; find the URL via
SearXNG `"<EN title>" preview pdf`). Used for item 12 (Evers-Fahey, 9781317219583).
### 15. Shop directory (HSE bookshelf34 pattern) — where RU psychoanalytic books are sold
- HSE «Книжный шкаф» (hse.ru/ma/therapy/bookshelfNN) = a DIRECTORY of shops, not a book list:
NikBook (nikbook.ru — WBS shop, `?q=` search is fuzzy/works), Когито-Центр, **Скифия**
(skifiabook.ru — publisher/shop, JS catalog, has a full price list download), marketplaces.
- NikBook Jungian finds (2026-07-09): Тоцци «Активное воображение в теории, практике и обучении»,
Конгер «Юнг и Райх. Тело как Тень», Эдингер «Библия и психе».
- beta2alpha.ru — RU publisher (Tilda site; sells via Ozon seller beta-2-alpha); Jungian stock TBC.
### 16. Publisher sites — RU full texts and bibliographies
- hi-human.org (Living Human Heritage RU) — publishes book forewords/sections as pages (source
of the Neumann EN-works list with RU titles).
- castalia.ru / castaliasilvasacra.ru (Клуб Касталия) — author articles + shop.
**Full catalog crawl-able (2026-09-24):** `/collection/all?page=N` — 48/page, 27 pages = 1285 товаров,
plain curl, заголовки в h3/h2/alt; кэш `data/catalogs/castalia-full-catalog-2026-09-24.md`
(в rutitles DB). **(PDF)-твины** (`/product/<slug>-pdf`) = прямые PDF-продажники = download-источник.
Product JSON-LD: name/sku(CAS+ISBN)/price; year/ISBN-блок «Характеристики» = JS (не читается).
- daimon-verlag.ch / chironpublications.com / spring-publications.com — EN Jungian presses
(spring-publications unreachable from this box; use Google Books ISBN instead).
### 17. alib.ru — RU/KZ book marketplace, 3.2M listings (used + new, incl. small press)
- Tool: `tools/alib.py "query"` (stem search! matches all inflections — unlike libgen).
- Endpoint: `https://www.alib.ru/find3.php4?tfind=<query>`; **query and pages are CP1251**
(utf8 → mojibake). Phrase mode: double quotes. Also: author-first, year range, ISBN digits,
price range. `>Купить<` links = listing count.
- Value: catches small-press / out-of-print editions shops don't carry (found: Furth 2nd ed.
2014, Turner handbook 2015, Neumann КДУ/Маниф/Питер editions). Seller titles can be wrong —
verify against RSL before trusting (e.g. «Челокес и миф» Kastaalia 2018 — not in castalia.ru
catalog → seller error).
- Also: alib.top (Ukraine), 33ob.ru (vinyl), amarka.ru (stamps) — sister sites.
### 18. Local CRW renderer (fastcrw.com, Firecrawl-compatible, localhost:3000)
- Tool: `tools/crw.py <url> [--html] [--wait N ms]` — POST /v1/scrape {"url","formats":["markdown"]}
- Use for JS-heavy sites curl can't read (soznanie.ast-academy.ru festival site rendered fine).
- Limits: Ozon (FAB challenge) and chitai-gorod SEARCH are API/anti-bot protected — CRW returns
nothing. chitai-gorod PRODUCT pages DO render via CRW (see source 9).
- Insales shops (castalia.ru) don't need it: product data is server-side JSON-LD (curl OK).
### 19. NLR / РНБ (National Library of Russia) via Primo (primo.nlr.ru) + CRW
- Tool: `tools/nlr.py "query"` — free-text search (all fields, words ANDed), page 1 (20 rows).
- Search is client-rendered → the tool renders the `search.do?fn=search&ct=search&
vl(freeText0)=QUERY&vid=07NLR_VU1&mode=Basic&initialSearch=true` URL via CRW and parses
title/author/year/holding (+ doc IDs when the linked layout renders).
- Full record (GOST description, incl. translators): CRW on
`display.do?tabs=detailsTab&ct=display&fn=search&doc=<07NLR_LMS#########>&displayMode=full&vid=07NLR_VU1`
(plain curl gives only the shell — details tab is XHR).
- Value: complements RSL — own St. Petersburg holdings + GOST records; author rows carry
life dates and NLR auth IDs. Count line parsing is flaky — trust the parsed row list.
- Verified 2026-07-09: all 35 section-04 surnames swept (see sections/04-pictures/SWEEP-R7.md).
### 20. RSL direct (aleph.rsl.ru, Ex Libris Aleph) — PARTIAL
- find-a flow works: GET /F/-?func=file&file_name=find-a → parse action URL (session token)
→ POST func=find-b&find_code=WAU&request=<name> → results 10/page with pagination.
- Full record: follow the result row's `full-set-set` link (session-dependent); gives
translator, series, ISBN, UDC (verified: Нойманн «Амур и Психея» МИФ 2024, ISBN
978-5-00214-519-5, пер. М. Виноградова).
- Quirk: find_code WTI (title) and SYS (ISBN) return EMPTY 0-byte responses; only WAU
(author) reliably works so far. libgen proxy (tools/rsl.py) still covers title/ISBN queries.
- search.rsl.ru — same catalog, Yii app, no JSON API found (2026-07-09).
### 21. ISBN validation pipeline — `tools/isbnval.py` (2026-07-15)
- Per ISBN: (1) libgen.vg biblioservice `type=isbn` — checksum + registered publisher group;
(2) Open Library `search.json?isbn=` — title/author/year; (3) Google Books `?vid=ISBN` curl
— title tag, 200/404; (4) NLR free-text for RU misses. Random pauses 2–5s, resumable.
- Verdicts: VALID (found w/ matching title) / STRUCT-ONLY (real ISBN, prefix registered,
not indexed — small-press RU) / BAD-STRUCT. Output TSV: `data/isbn-validation.tsv`.
- Found in sec04: 2 bad-check-digit typos (5-89613-003-6, 5-263-00368-2), 2 publisher-prefix
mismatches (Ленанд vs URSS; Академический проект vs Азбука-Аттикус), 2 co-publishing
prefixes (Петроглиф for ЦГИ 2020; Пальмира for Касталия 2025).
- **triumph.ru** (РГБ ISBN search, user-provided): `https://www.triumph.ru/html/serv/find-isbn.php?isbn=<13digits>`
(relative to the poisk-isbn.html page!) → returns links to РГБ/МГУ/РНБ/БЕН РАН/ГПНТБ catalogs
pre-filled with the ISBN. Itself no data; МГУ (nbmgu.ru) results are JS-rendered, БЕН РАН
(Koha) is scrapable but small collection. NLR absence ≠ ISBN invalid.
- **isbnsearch.org**: nice per-ISBN pages (author/publisher/year) but rate-walls after ~10
requests ("Please Verify to Continue", no recovery in 7 min) — not for batches.
### 22. MGU Scientific Library (nbmgu.ru) — server-rendered, curl-able (2026-07-15)
- Tool: `tools/mgu.py "query" [FIELD] [pages] [method]` or `--multi "AUT:X" "TIT:Y"`.
The "JS" advanced search is a plain GET: `/search/?adv=1&q1=<q>&f1=<F>&v1=<M>&cat=BOOK[&p=N]`.
- Fields: ANY/AUT/COA/TIT/KEY/RUB/YEA/PLA/PUB/SER/ISB/ISS/NBM; rows AND between q1..q3.
Method v: 0=Слова (stemming, default) · 1=Словосочетание · 2=Начинается с · 3=Дословно.
- "Всего: N" + GOST row (title/authors/notes) + "City : Publisher, Year" + shelf code + uid link.
20 rows/page, p=0-based; **out-of-range p silently falls back to page 0** (dedupe by uid!).
- **ISB field NOT populated** (0 hits, 3 formats × control ISBNs) — no ISBN search; use
AUT/TIT/SER/PUB. `storing.aspx?uid=` page = holdings/order only (no full record).
- Value: second RU catalog (complements RSL/РНБ), strong **SER** sweep (e.g. «Библиотека
аналитической психологии» → full series list in one query), GOST rows w/ translators.
- Anchor titles in result rows contain a raw `>` (title="Хранение<br/>Заказ") — regexes must
not use [^>]+ across the anchor attrs.
### 23. bookvoed.ru — curl-able, ISBN in page title (2026-09-20, user-provided)
- `https://www.bookvoed.ru/product/<slug>-<id>`: HTTP 200 with plain curl; `<title>` contains
`Название (Автор) (ISBN)` — a fast existence + ISBN check. Full specs in HTML.
### 24. litres.ru book pages — curl-able JSON-LD (2026-09-20, user-provided)
- `https://www.litres.ru/book/<cat>/<slug>-<id>/chitat-onlayn/`: HTTP 200 with curl; JSON-LD has
name/author/publisher (translator/ISBN/pages often absent). «Кратко» (Culture-Multur abridgments)
are distinct ISBNs from the originals — don't conflate. Good for title-variant discovery.
### 25. royallib.com / pda.coollib.in — free RU full text, curl-able (2026-09-20, user-provided)
- `royallib.com/book/<author>/<title-slug>.html` — 200 with curl, annotation + text pages;
`pda.coollib.in/b/<id>-<slug>/readp?p=N` — page-by-page text. Use for CONTENT verification
(e.g. Eliade «Тайные общества» = Rites and Symbols of Initiation: preface = Haskell Lectures
1956, University of Chicago — fingerprint of the EN original). Both = download supplements.
### 26. book.ivran.ru + opac.liart.ru — small RU library OPACs, curl-able GOST records (2026-09-20)
- `book.ivran.ru/book?id=<N>` — «Книги сотрудников ИВ РАН» DB: full GOST rows (title, editor,
city/year/pages, ISBN). Find via SearXNG `site:book.ivran.ru`.
- **opac.liart.ru (РГБИ) — platform switched to OPAC-Global (DITM-Global)** (2026-09-21, probed):
- `/opacg/` guest session: POST `arg0=GUEST&arg1=GUESTE&TypeAccess=PayAccess` to
`/cgiopac/opacg/opac.exe`; session `9096453`/`GUEST` (hardcoded in JS).
- Search = POST `/cgiopac/opacg/direct.exe` FLAT fields: `_service=opacfindd.FindView`
(or `FindSize`), `query/body="<label> <value>"`, `iddb=<N>`, `userId`, `session`,
`_xsl=_version=2.7.0/1.1.0` — full protocol in `tools/liart.py` docstring.
- **BLOCKER:** valid search labels + iddb come from server InfoDB (admin-only for GUEST;
iddb 1–4 absent, 7+ exist). Record pages `/record/<db>/<hex>` ARE curl-able (GOST in <title>).
SearXNG `site:opac.liart.ru` index is sparse (author queries 0, `site:.../record` 20).
- **Tool: `tools/liart.py`** — default mode = SearXNG record-URL fallback (works today);
`--direct` mode ready, needs `data/liart-params.json` filled from the browser
(Поиск → Базовый → F12: `input id=basic_<label>` / `select id=field_1`). TODO remains.
### 27. maap.pro «Библиотека» — RU article translations of Jungian authors (2026-09-20, user-provided)
- `https://www.maap.pro/biblioteka/stati/` — МОСКОВСКАЯ АССОЦИАЦИЯ АНАЛИТИЧЕСКОЙ ПСИХОЛОГИИ
publishes **article-length** RU translations (e.g. Stevens «От Юнга» = fragment of *On Jung*,
author bio lists the standard RU titles: «О Юнге» 1991, «Нить Ариадны: путеводитель по
символам человечества» 1999, «Двухмиллионо-летняя Самость» 1993). Value: (1) essay-level 🔶
status for ❌ books; (2) canonical RU title naming from the author bio.
### 28. triquadratapubl.ru (Три Квадрата) — publisher site, curl-able product pages (2026-09-20)
- `triquadratapubl.ru/product/<slug>/` — WPCOM product pages with full specs (publisher,
ISBN, series, annotation). Kerényi «Мифология» (2012, 504 pp., ISBN 978-5-94607-150-5/
-260-1, ред. Клименко А., сост. И. Надь) = «Сборник статей по античной мифологии» — a
compilation, NOT *Essays on a Science of Mythology* (item 03 stays ❌).
### 29. fantlab.ru — structured work pages for RU translations (added 2026-09-22, sec05)
- `fantlab.ru/work<id>` — server-side HTML, curl-able (CRW for the publications block). Blocks:
«Перевод на русский» (translator + year + N изд.), «Издания» (period/language/translator filters),
«Входит в» (anthologies — e.g. «Бригантина 72-73»). Find work id via SearXNG
`site:fantlab.ru "<RU title>"`.
- Value: **Soviet-era / adventure / ethnographic / sci-fi books** that RSL and libgen miss
(found: van der Post «Потерянный мир Калахари» 1973, пер. Л. Деревянкина — sec05 #18;
Bjerrre «Затерянный мир Калахари» Мысль 1964 = homonym trap). Use for pre-1991 items.
### 30. predanie.ru + livejournal — older-edition pages & community oracles (added 2026-09-22, sec05)
- `predanie.ru/book/<id>-<slug>/` — curl-able book pages for older RU editions (Ладомир etc.),
GOST-like specs. Find via SearXNG `site:predanie.ru "<RU title>"`.
- **livejournal.com** reviews («Книжная полка культуролога» etc.) — CRW-rendered; often cite
BOTH RU title variants of one book (Benedict «Модели культуры»/«Образцы культуры»). Good
oracle for title-variant discovery; not a bibliographic source.
- **jungnewyork.com** — EN publisher/institute pages (per-book: TOC, praise, EN metadata);
NOT a RU source (book_mi.shtml = Adams EN page).
### 31. psyinst.moscow «Библиотека» — Институт психоанализа (Москва): полные RU-тексты (added 2026-09-26, sec08)
- `psyinst.moscow/biblioteka/?part=article&id=<N>` — полные тексты RU-изданий психоаналитической
литературы (OCR'd, встроены в HTML страницы). Найдено: Шпиц «Психоанализ раннего детского
возраста» (Per Se/Унив. кн. 2001, пер. Боковикова/Старовойтова, 159 pp) = *Psychoanalysis of
Early Childhood* (1957) — id=1045 (user-подсказка, 2026-09-26).
- Value: (1) full text для books, отсутствующих на libgen/flib (RU-даунлоуд-резерв);
(2) fingerprint-верификация RU издания по тексту. Искать: SearXNG `site:psyinst.moscow "<RU title>"`.
## Lessons from sec05 (2026-09-22) — what made the 7 rescues possible / what to do next time
1. **User oracle table = highest-ROI step.** Initial ❌ list (19) was 37% wrong; user checked the
RU-title-guess table against OZON/labirint memory and rescued 7. **Sec09 подтверждение (2026-09-27):
user дал прямые URL-ы (OZON + free-PDF-зеркала lib.uni-dubna.ru / solivint.ru) = 3/11 rescues (27%);
sec10: 2/27 (7%) — Jamison «Беспокойный ум» (Альпина 2017) + Saks «Не держит сердцевина».
Free-PDF-зеркала (university library sites, solivint, psyliterature.narod.ru) = недооценённый
download-канал для старых RU-изданий.** **Convention: at section end,
present the ❌ table (author, EN title, guessed RU author/title, channels probed) to the user**
before finalizing MISSING.md.
2. **RU titles are NOT calques** — the actual titles found: «Очерки сравнительного РЕЛИГИОВЕДЕНИЯ»
(not «религии»), «Аспекты мифа» (translated from FRENCH *Aspects du mythe*, not from EN *Myth and
Reality*), «Модели культуры» (not «Узоры/Образцы культуры»), «Потерянный мир Калахари» (the
«Затерянный» variant is a DIFFERENT book — Bjerrre). Always probe 2–3 RU title variants;
for authors with FR/DE originals (Eliade, van Gennep) also try the FR/DE-based RU title.
3. **rutitles diff MANDATORY before ❌:** sec05 пропустил его и потерял бы 3/7 rescues
(все были в data/ru-titles.jsonl). `tools/rutitles.py both "<EN title>"` по каждому item
(в negatives gate ниже).
4. **flib authorall for prolific RU authors** (Элиаде a/2943 = 49 books) caught «Тайные общества»
(sec05 #13) that libgen title search missed. Always run flib author sweep for authors with
>5 RU books.
5. **libgen mislinked-file trap + KEY vs f_id (2026-09-25, уточнение):**
(a) Edition files can be mislinked (141685581 «На пороге инициации» carried an ePubLibre
Spanish Lermontov file; the sec05 «pdf» batch carried English academic papers; the 4 van Gennep
epubs were Spanish fiction). Sizes pass lgdl's md5 check. **Mandatory: after every libgen
download, content-verify: `pdfinfo` page count vs RSL record + first-page text
(pdftotext -l 3 / fb2 body head / epub xhtml head) must contain the RU/EN title/author.**
For recent books prefer flib (clean files); flib `/b/<id>/download` → 302 → direct
`static.flibusta.is` URL.
(b) **KEY vs f_id (системная ловушка, наш баг, не libgen):** в JSON-карточке издания
`files` = `{public_key: {f_id: X, md5: Y}}` — КЕЙ и f_id-ЗНАЧЕНИЕ = РАЗНЫЕ файловые записи.
`object=f&ids=<key>` возвращает устаревший/чужой файл (Solomon: key 92709688 = 20MB рус.
антология; f_id 93253044 = настоящая 2MB-книга).
**Правило: качать и пречекать ТОЛЬКО по f_id-значению из files-мапы (md5 для контроля)** —
утренний «24 mislinked EN» 2026-09-25 = качали по key; по f_id — rescues sec06.
Файл-уровневый поиск (`objects[]=f` в lg.py) находит файлы, не привязанные к «нашему» изданию.
6. **Collection-inclusion questions** (#27/#28): download the smallest fb2, parse the TOC
(`<section name="title">`), grep for the EN book title + RU calque. «Символ и ритуал» (Наука
1983) CONTAINS «Ритуальный процесс» (full text, core section) but NOT «Лес символов»
(citations only).
7. **flib pdf download**: `/b/<id>/pdf` may return an HTML stub; `/b/<id>/download` → 302 →
direct `static.flibusta.is:443/b.usr/<File>.<ext>` URL — curl that directly.
8. **Content-verify gate (MANDATORY):** после КАЖДОГО даунлоуда — magic bytes + `pdfinfo` pages
+ first-page/head text = ожидаемый title/author; для статей — автор+журнал+страницы = RAW и
это статья, НЕ рецензия. Прошёл = discard + re-source (flib/lg другой ed). **Pre-download:
`tools/lgdl.py precheck <f_id> <edition_id>`** (0 OK / 1 SUSPECT / 2 BAD; `dl … <edition_id>`
запускает precheck автоматически). Mislink = next f_id → next edition → flib; отклонённые
f_id в MANIFEST notes. (Логика precheck v2: reverse-map + locator, см. докстринг скрипта.)
9. **EN-queue infra (переиспользуемый паттерн):** `queue.sh` = per-file
`[ -f out ] && skip || timeout 1500 python3 -u lgdl.py dl ... >> "$LOG" 2>&1` + **watchdog**
(рестарт если очередь умерла без DONE-маркера; skip-existing делает рестарт безопасным) +
`setsid nohup`. Ловушки: (a) `bash -n` на сгенерированный скрипт ОБЯЗАТЕЛЕН (one-liner
`dl() { … }` без `;` перед `}` = тихая смерть);
(b) python stdout блок-буферизуется в файл — всегда `-u`;
(c) `ads.php` throttle под sustained load — пауза 30–60 мин, НЕ долбить;
(d) NEVER `pkill -f lgdl.py` (убьёт и саму очередь) — kill worker по comm;
(e) dual-queue corrupt cookie jar. AA (annas-archive.gl) = 403 anon отсюда; i2p not worth it.
## Content audit (method; ran on sec01–07, 2026-09-25/26)
Tool: `tools/audit_content.py <sec>...` (samples first/mid pages of every file, scores title
overlap, flags SUSPECT/SMALL/NO-TEXT/NO-CYR). Run at section end. Триажа флагов:
1. **Mislinked file** (файл = другая книга/рецензия/фрагмент) → удалить, re-source (flib/lg),
записать в MANIFEST Notes. Превью/фрагменты (2–14 pp, chapter previews) = не издание.
2. **Foreign OCR layer** (файл OK, текстовый слой на чужом языке/мусор) → keep + пометка,
чистая twin-копия в приоритете.
3. **False positives (не чинить):** windows-1251/UTF-16/KOI8 fb2/txt (декодирование, не
порча), image-сканы без OCR (проверить рендером: pdftoppm/ddjvu), spaced-out OCR titles
(«T H C T F»), library stamps, VeryPDF-водяные знаки.
4. **Bonus rule:** related-but-different книга = bonus-файл с пометкой `Бонус: ДРУГАЯ книга
(...)`, status не менять; переименовать файл, если false lang/year в имени.
5. Контент-верификация после КАЖДОГО даунлоуда — см. Lessons (gate).
## Catalog caching convention
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source
URL + date in the header) so negative verification stays re-runnable without re-crawling.
Examples: `cogito-yungianskaya-series-2026-07-17.md` (30 pp. CRW crawl, 518 titles),
`litres-140409-2026-07-17.md`, `opp-misp-reading-list-2026-07-17.md`.
## Pipeline v3 (2026-07-17) — RU-title candidates & batch sweeps
- **`tools/rutitles.py`** — the "diff, not search" tool:
- `build` — rebuild `data/ru-titles.jsonl` from `data/catalogs/*` + preserve flibusta entries
- `addflib <authorId>` — append flibusta.is authorall titles to the DB
- `translate "EN Title"` — RU token bag via `data/ru-dict.tsv` (540+ EN→RU pairs, phrases first,
observed Kogito/Castalia/Peter conventions) + list of UNTRANSLATED words (finish manually)
- `match "RU tokens" [N]` / `both "EN Title"` — Jaccard over **pymorphy3-lemmatized** tokens
(case endings handled: «великой матери» → «Великая мать» 1.25; order-agnostic)
- DB (2026-09-24): 1977 titles (Когито-518 + ЛитРес-24 + ОПП/МИСП-55 + **castalia-full-catalog-1285**
+ flib a/5272, 57639, 118921, 122883, 193690, 193689); `build` пересобирает из data/catalogs/*
(новый каталог = просто файл в data/catalogs/, строки «название [таб url]», #-строки игнор)
- **Limits (honest):** generic RU titles (Wolff «Введение в основы…») invisible — publisher
sweep (Phase 2b) stays a separate phase. Translation ≠ identity: verify real hits.
- **pymorphy3** installed (user-site): lemmatization for RU token matching.
⚠ pymorphy2 is BROKEN on Python 3.12 (`inspect.getargspec` removed) — use pymorphy3 only.
- **`tools/oa.py`** — OpenAlex/Crossref (Phase 0, see source 10b).
- **flibusta .su dropped** from the search pipeline (mirror of .is, reduced functionality).
- **`tools/sweep.py`** — batch driver for Phases 1–2 (per author/item: rsl + lg + alib + flib +
cogito + SearXNG-OZON-snippet parse + rutitles diff; optional `--nlr` for НРБ via CRW, slow).
Query log: `data/queries.log` (JSONL). Sweep output: `data/sweeps/secNN/<nn>-<slug>.md`.
## Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)
- bookmate.com (403 bot protection), gtmarket.ru (JS app, no HTML content)
- other libgen mirrors (libgen.lol → error page; .vg works with browser UA)
## Identity check protocol (MANDATORY before trusting a hit)
1. **Life dates** in RSL author field (e.g. `Нойманн, Эрих [1905-1960]`).
2. **EN↔RU cross-mapping:** the found person's *other* RU books must map to known EN titles
(e.g. found «Страх феминного» + «Любовь и Душа» → EN *Feminine* + *Love and its Opposites*
→ same Erich Neumann).
3. **Co-publisher signature:** known translation pipelines (Vakler/Рефл-бук 1990s,
Касталия, Прайм-Еврознак, АСТ/ЭКСМО, Питер «Мастера психологии»).
4. **Content spot-check** of the actual target title (open edition page / TOC) when equating
an RU title with an EN title that sounds different
(e.g. *The Archetypal World of Henry Moore* ⇄ «Искусство и творческое бессознательное» — verify it is about Moore).
5. **Homonym rejection:** different first name / field / era → reject or mark ambiguous
(trap found: Верена Каст ≠ Вернер Каст; «Каст» search returns Verena Kast books).
## Workflow (per section) — v3 (2026-09-24: +Phase 1.5 batch catalog sweep, +канон. RU имя)
**Phase 0 — Full names + identity** (`authors/`):
1. Resolve every initial-only author to a full EN name: `tools/ol.py 'title:"EN Title"'`
(author_name) → else `tools/wd.py` (QID + RU label) → else Google Books ISBN page →
else publisher/review page. Write/extend the **dossier** `authors/<slug>.md` (EN full
name, dates, RU name to query, name source); then add a one-line entry in the thin
index `sections/<n>/AUTHORS-EN.md` linking to it. If the dossier already exists from
another section, just append the new list item + link it.
2. For big names: `wd.py` gives life dates + the OFFICIAL RU label (best RSL spelling).
**`tools/rulabel.py <QID>` = каноническая кириллица (ru label + aliases)** — всегда брать
до свипов: Циммер ≠ Зиммер, Корбен ≠ Корбин, Шолем ≠ Шулем, Клюгер ≠ КлуGER.
3. Build the transliteration matrix per author (2–3 variants: Нойманн/Нойман, Кох/Коч,
Калфф/Кальфф, Ферс/Фёрт) — run every query with ALL variants. **Каноническое RU имя
(rulabel/досье) — ПЕРВЫЙ вариант матрицы; стандартная транслитерация — только запасная.**
**Phase 1 — Author sweep (EN→RU):** per author: `rsl.py "<RU First> <RU Last>"` (ВСЕГДА с
каноническим RU именем из досье/rulabel, не с транслитерацией наугад) → `lg.py "<RU Last>"`
→ `cogito q=<RU Last>` → `flibusta authors "<RU name>"` → livelib/chitai-gorod author page.
Apply identity check. Fill author dossier.
**Phase 1.5 — Batch catalog sweep (2026-09-24, после Phase 0–1, ПЕРЕД карточками):**
один проход по ВСЕМ item секции — до написания карточек. Определяет «карту покрытия»:
какие книги точно в RU/EN каталогах (✅-кандидаты), какие нет. (Опыт sec06: #08 Никлаус
из Флюэ пойман именно здесь, а не per-card.) Шаги:
1. **Кросс-свип по обработанным секциям (PRIORITY 1):** grep всех карточек `sections/0N-*/
*.md` (H1 + filename) по каждому item — книги из прошлых секций не ищем заново, а делаем
cross-ref-карточку (механика: python-скрипт по списку ключевых слов item; пример sec06:
6 кросс-рефов 01/02/03/05/06/11 за один проход). Досье авторов тоже: «Verified RU works»
может уже содержать ответ.
2. **`tools/rutitles.py both "<EN title>"` по каждому item** (DB = Касталия-1285 + Когито-518
+ ЛитРес + ОПП/МИСП + flib-авторские списки); подозрительные ✅-кандидаты — открыть продукт
(castalia/esxatos/cogito) и взять изд/год/ISBN.
3. **Каталоги издательств:** castalia-full-catalog (rule выше) + flib **authorall** по авторам
с >5 RU-книгами (фон Франц a/57639, Элиаде a/2943, …).
4. **EN-свип:** `oa.py` по EN-названиям (OpenAlex/Crossref — ISBN/полные имена; фаза 0,5);
soft libgen queries (`<author> <1–2 title words>`, см. «EN original sweep» ниже) +
`ia.py search` — batch, до даунлоуд-фазы, не после.
Результат фиксируется в INDEX.md (статусы до Phase 2). Затем Phase 2 добирает конкретные
книги, gate v3 — только финальная галочка.
**Phase 2 — Title-gap sweep** for still-unmatched books: `rsl.py "<RU title words>"` (check
author field!), `lg.py` 2–3 whole-word title combos, cogito, flibusta by RU title, GBS SPb via
SearXNG `site:gbs.spb.ru`.
**Phase 2b — Publisher sweep (RU→EN, reverse direction):** per Jungian RU publisher/imprint
(Касталия, hi-human/Living Human Heritage, ЦГИ/Т8, Когито, Фантом Пресс, Азбука, Серебряные
нити, Питер, Ваклер, Добросвет, Деметра, НикБук, Скифия, beta2alpha, Эксмо/АСТ психология):
enumerate the full Jungian catalog (site search or RSL **series** queries: «Юнгианская
психология», «Библиотека аналитической психологии», «Классики зарубежной психологии») → map
every title to an EN original → backfill both missing items and future sections. This catches
books author queries can never hit (generic titles like «Сборник статей», reworked titles).
**Phase 2c — Community reading lists (oracles):** ISAP Zurich, HSE (hse.ru/ma/therapy),
МИСП, Jung Institute Moscow, education-psy.ru sandplay list, carljung.ru, koob.ru — these
explicitly state which books have RU editions (or lack them). Scrape each, diff vs our ❌ list.
**Phase 3 — Content verification (for suspicious candidates):** when an RU title's mapping to
an EN title is not obvious, verify 1:1 against the EN original: archive.org OCR text (TOC /
editorial note) vs RU essay/chapter list (koob/livelib/preview). Example: Neumann «Сборник
статей» = Art and the Creative Unconscious: Four Essays (4/4 essay match).
**Phase 4 — Edition collection:** ALL editions from ALL sources (title variants!), full metadata
(publisher/year/pages/translator/ISBN) — prefer RSL/GBS SPb records (GOST) over shop data.
**Phase 5 — Book notes + commits** (one commit per completed book or meaningful author gain).
When a section is complete, its `SUMMARY.md` must contain a consolidated
**"Publication status at a glance"** table (one row per reading-list item) with these columns:
`# | EN book | EN ISBN | RU book | RU ISBN | DL EN | DL RU`. This is the single
quick-reference view — it records not just ISBNs but, for each item, whether the EN original
and the RU translation are actually **downloaded** (file count + format, e.g. "5 pdf", "2 fb2",
"1 djvu", "1 doc", or "sec04: N pdf" for cross-section items, or "— (restricted)" / "— (absent)" when
not downloadable). Legend for ISBN columns: **✅** = ISBN found; **⚠** = published but no ISBN
(not stated / new ed. / OZON-prefix only / truncated); **❌** = no such edition (verified absent);
**—** = the book itself has no such edition (essay / Standard Ed. vol / pre-ISBN). End the table
with a one-line totals row (DL EN / DL RU file counts). Cross-check the DL columns against the
section's `downloads/<nn>-*/MANIFEST.md` so the two never disagree.
Negatives gate (before ❌) — **v3 (2026-09-24, after sec06 rescues)**:
author sweep (RSL + libgen + cogito + flibusta **authorall** for prolific RU authors) +
**`tools/rutitles.py both "<EN title>"` diff (MANDATORY; DB includes castalia-full-catalog)** +
**castalia.ru catalog check (MANDATORY for Jungian items — thematically fitting publisher; see rule below)** +
2–3 RU title variants
(calque + SearXNG/web_search `"<EN title>" перевод/на русском` + FR/DE-original variant for
Eliade/van Gennep-type authors) + livelib author page + one shop/retail check (cogito person
page for identity) + **fantlab.ru for pre-1991/adventure/ethnographic** + predanie.ru for older
editions + **user oracle table at section end** (see Lessons). Log every RU query in MISSING.md.
Before downloading from libgen: verify file locator/title (mislinked-file trap).
**Castalia catalog rule (2026-09-24, user instruction):** Клуб Касталия — тематически
подходящее издательство (Юнг/аналитическая психология), его каталог = ОБЯЗАТЕЛЬНЫЙ этап поиска
для каждого item (не только ❌-гейта): (1) `data/catalogs/castalia-full-catalog-2026-09-24.md`
(1285 товаров) уже в rutitles DB — `rutitles.py both` всегда ищет и по нему; (2) при подозрительном
❌ (юнгианская тема) — прямое grep каталога по 2–3 RU-вариантам названия; (3) кэш РЕФРЕШИТЬ
(перекроулить `/collection/all?page=1..N`) если он старше ~6 мес или перед финальным ❌ секции;
(4) (PDF)-твины товаров = download-источник (см. source 16).
Special cases:
- **Jung's CW volumes:** translated and widely available — don't deep-search; note the RU
volume for the item. Факты: АСТ сер. «Философия – Neoclassic» публикует CW последовательно
(CW3 = «Психогенез душевных болезней» 2025, пер. Чечина, 978-5-17-139315-1; CW2 пока нет —
ре-проб). CW15 = «Дух в человеке, искусстве и литературе» (Харвест 2003); Red Book =
«Красная книга (Liber Novus)» (Касталия 2025, пер. О. Комков). Калка-ловушка: «Психогенез
**душевных**» ≠ «…психических».
- **Essays/chapters** (e.g. Jaffé in *Man and his Symbols*): mark 🔶 if only the parent
book/anthology is available in RU.
## Politeness
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.
## EN original sweep (soft libgen queries) — method (validated 2026-09-20, reused sec05/06/07)
Exact full-title libgen queries miss a lot (indexed titles vary: reprints, e-books, variants).
The validated approach for closing the EN download gap of a section:
1. **Build the gap list** from `downloads/0N-*/MANIFEST.md` («Not downloadable» block) +
SUMMARY «at a glance» DL EN column (rows with «— (restricted)» / «—»).
2. **Soft libgen queries, 2–3 rounds** (author-centric, not title-guessing):
- round 1: `<author surname> <1–2 distinctive title words>`;
- round 2 (0-hit items): `<author surname>` alone, filter results by title/first name;
- round 3: alternate wording (shortened title, co-author, publisher word).
Whole-word matching still applies (full words, case-insensitive). `tools/lg.py` prints
`(unparsed…)` on 0-result pages — treat as 0 but confirm once with raw curl grep `edition.php?id=`.
3. **Collect candidates → batch metadata:** `https://libgen.vg/json.php?object=e&addkeys=*&ids=<comma-list>`
(≤10 ids) → `files` dict gives `f_id`+`md5` per file; then `object=f&addkeys=*&ids=<f_ids>`
for extension/filesize. Filter by year/edition to match the reading-list book (identity check:
title + author + era; e.g. SE15 vs SE22 look similar in truncated titles).
4. **Download queue (serial):** `tools/lgdl.py dl <f_id> <outdir> <name>`; resume works via `.part`
(curl `-C -`); the tool retries with fresh get.php keys. Huge files (>100 MB) can crawl at
~10 KB/s on libgen CDN (random stuck offsets) — let it stall-detect, or abandon in favor of
another edition of the same book (e.g. Great Mother: 643 MB Princeton 2015 abandoned →
64 MB Bollingen 1991).
5. **Monitor-kill trap (2026-09-20):** NEVER `pkill -f "lgdl.py"` (or any pattern that appears in
the queue bash's own command line) — it kills the queue itself. Kill only the worker by comm:
`ps -eo pid,comm,cmd | awk '$2=="python3" && /tools\/lgdl/ {print $1}' | xargs -r kill`.
6. **Verify after:** magic bytes (`%PDF` / `PK\x03\x04` for epub; discard HTML stubs) +
`pdfinfo` page count for small files — 2–14 pp «pdf» are fragments/essays, keep but mark
«фрагмент» in MANIFEST (e.g. Way of All Women 1933 = 2-pp, Self in Transformation = 14-pp JAP essay).
7. **archive.org (ia) as complement** for trade-restricted items: `tools/ia.py search "…"` →
open items only (`RESTRICTED` = borrow-only). Jung CW: single open item `CarlJungCollectedWorks`
(vols 1–18 + seminars, PDF+EPUB, direct `/download/<item>/<file>`).
8. **Docs to update after:** SUMMARY at-a-glance DL EN + totals row, MANIFEST (per-item rows with
lg f_id / ia item; «Not downloadable»), per-card Downloads table + Notes line.
One commit per section. (Результаты прошлых секций — в их MANIFEST/Summary + git log.)
## Dedup rules (user, 2026-09-25) — КАЖДАЯ КНИГА ИЩЕТСЯ/СКАЧИВАЕТСЯ ОДИН РАЗ
- **Book level:** `data/MASTER-LIST.md` (генерируется `tools/xref.py --write`) — все 12 секций,
каноническая карточка + «Also listed in». Триггеры: (1) разово уже сделано 2026-09-25 (карта
для scope-решения 07–12); (2) **в начале каждой новой секции** — `python3 tools/xref.py --write`
и работать по её NEW/xref-меткам (raw↔raw-метки уточняются против карточек пред. секции);
(3) **в конце секции** (после даунлоудов) — plain `python3 tools/xref.py` = stale-line-чек
манифестов (см. правило ниже). RAW новой секции кладётся в sections/0N/RAW.md (или
data/isap-raw/), xref.py подхватывает автоматически. Скрипт локальный, ~1 c.
- **File level:** файл хранится ТОЛЬКО в первой секции, где книга появилась; поздние секции —
cross-ref «secNN: <file>» в MANIFEST/карточке. (2026-09-25: 17 md5-дублей удалено
`tools/dedup_files.py`.)
- **Профилактика дублей:** перед скачиванием в новой секции — md5-scan downloads/ по книге
(или по MASTER-LIST: если каноническая секция уже скачала — не скачивать).
- **MANIFEST «Not downloadable»** — не источник истины по кросс-секционным книгам:
перед ❌/gap-выводами прогонять `tools/xref.py` (раздел stale-lines) — phantom-строки
(Aion в sec05) и устаревшие «not found» после rescues в других секциях.
**Семантика (user, 2026-09-25): секция = лог ИСТИННЫХ дыр. Дыра закрыта = СТРОКУ УДАЛИТЬ**
(не помечать ✅ на месте, не вести «Rescued»-подраздел) — актуальное состояние в per-item
таблицах, история rescues в git log. Частичный rescue (EN нашлось, RU платное) = строка
сужается до оставшегося языка.
## Download phase conventions (user rule, 2026-07-18)
- **EN originals: libgen.vg FIRST** (the EN collection is bigger than the RU one — most trade
academic books: Routledge/Karnac/SUNY/Princeton/etc.), **archive.org second** (CW/Philemon
open scans, content-verification OCR, items libgen lacks). Flibusta for EN: tried
2026-07-18 — OPDS search is effectively RU-only (EN titles 0, author search returns only
Cyrillic pages); a one-shot exact-title probe is cheap, don't plan on it.
- **RU translations: libgen.vg + flibusta.is** (flib = fb2/epub when libgen has only scan-pdf or nothing).
- Per language, format priority: ALL pdfs if any pdf exists; else all fb2; else all epub; else all doc.
- File naming `NN-<slug>-<lang>-<edition>.<ext>`; **downloads в git LFS** (.gitattributes per sec),
MANIFEST.md — обычный текст.
## Lessons from sec08 (2026-09-26)
1. **pkill -f self-match (расширение ловушки 2026-09-20):** `pkill -f "en-queue.sh"` убил СОБСТВЕННУЮ
оболочку — паттерн оказался в command line bash'а, исполнявшего heredoc с этим же текстом.
Правило: pkill -f только по паттернам, которых НЕТ в текущей команде (или kill по pid из
`ps -eo pid,comm,args | awk '$2=="python3" && /lgdl/'`).
2. **EN-кью против ads.php-троттлинга:** паттерн, выживший в sec08: `en-queue.sh` (per-file
skip-if-exists + `timeout 1500` + tries=3 через `python3 -c "...lgdl.download(..., tries=3)"`)
+ `en-rounds.sh` (раунд → 25-min cooldown → раунд, до remaining=0). Один раунд = 10–20 файлов
до троттлинга; с tries=16 проклятый файл жрал 45 мин вместо 2.
3. **`ls a*.pdf b*.epub` возвращает non-zero, если ХОТЯ БЫ ОДИН паттерн не матчит** (даже если
другой сматчился) — счётчики «remaining» в таких if'ах завышаются. Проверять через
`ls pattern 2>/dev/null | grep -q .` или по каждому расширению отдельно.
4. **Libgen-файлы-«книги», которые на деле не книги (6 отклонений за секцию):** 2–3-стр. BOOK
REVIEWS (08, 14, 18, 38), 1-стр. «Book Reviews» фрагмент (30), 24–26-стр. JAPA presentation/
article вместо книги (33). Precheck не ловит (locator совпадает по теме!) — ТОЛЬКО
post-download: pdfinfo Pages (<10 pp = подозрительно) + first-page text.
5. **ia-ресурсы секции:** CarlJungCollectedWorks (CW4) и freud-complete-works (SE 24 тома, 5109 pp,
оба эссе 19/20 верифицированы в djvu.txt) — оба OPEN, прямые /download/ работают.
6. **DJVU на этой машине:** ddjvu = GUI (текстовый вывод = мусор), djvu2txt/djvu2png не установлены,
python djvutxt нет. Content-verify djvu = precheck (reverse-map + locator) — приемлемо при
MATCHES; при SUSPECT — поставить djvulibre-utils.
7. **Linter (check-md.py) fname_re** в tools/migrate_cards_v3.py — расширенный: +rtf (2026-09-26).