Init: research structure, tools (lg.py, rsl.py, sx.sh), AGENTS.md, CONTEXT.md, section 04 skeleton (39 book notes)

This commit is contained in:
Dmitry Kokorin 2026-09-08 22:10:07 +03:00
commit 9b8cf58a4d
55 changed files with 1421 additions and 0 deletions

123
AGENTS.md Normal file
View file

@ -0,0 +1,123 @@
# Jung Reading List → Russian Editions: Research Notes
## Goal
For every book in the ISAP Zurich reading list (https://isapzurich.com/en/library/reading-list),
find the **Russian edition(s)**: correct Russian title, publisher, year, translator, ISBN,
libgen download page (if exists), and РГБ (RSL) bibliographic record (if exists).
We work **section by section**. Current: **Section 04 — Pictures** (`sections/04-pictures/`).
Sections 1–3, 5–7 are skeletons for later.
## Repository layout
```
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
tools/ # search helpers (lg.py, rsl.py, sx.sh)
authors/ # one dossier per author (created on demand)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
```
## Conventions
- **One file per book**, named `NN-<slug>.md` (NN = order in the reading list).
- **Status marker** at top of each book file:
- `⬜ not searched yet`
- `🔎 searching in progress`
- `✅ RU edition(s) found`
- `🔶 partial` (e.g. only an essay/chapter available in RU)
- `❌ no RU edition found (verified)`
- **Book file contents:** human-readable editions table + raw **RSL JSON** block +
libgen edition URLs (`https://libgen.vg/edition.php?id=...` — contains download links;
direct file links rotate, don't store them) + notes (identity check, RU↔EN title pairing source).
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) hold: EN identity
(3–5 flagship EN titles, dates), RU surname candidates (all plausible transliterations,
German names often have several: Нойманн/Нойман/Нейманн), verified RU works with EN mapping,
homonym warnings.
- **Commit per completed book** (or when meaningful author info is gained), message like:
`sec04: <book> — <finding>`.
- Update `CONTEXT.md` for anything useful that fits no category.
## Sources and their status (verified 2026)
### 1. РГБ (Russian State Library) via libgen.vg biblioservice — PRIMARY for existence + editions
- Tool: `tools/rsl.py "Full Name"` (JSON, polite, 3s between pages)
- Raw URL: `https://libgen.vg/biblioservice.php?value=<query>&type=rsl&format=json`
- Rich records: title parts, **author with life dates** `[1905-1960]` (great identity anchor),
translator, publisher, city, year, pages, ISBN, series, UDC tags, even TOC field.
- **Limitations:** no pagination (fixed first page) → always query by **full first+last name**
for people, or 2–3 distinctive words. Loose matching (searches author AND title fields).
- Other `type=` values on this mirror are **dead**: worldcat, googlebooks, isbndb, udc
(all return "Nothing found" even for control queries). `type=isbn` only decodes ISBN structure.
### 2. libgen.vg full-text search — PRIMARY for readable editions
- Tool: `tools/lg.py "query"` (whole words only, res=100, prints pagination hint)
- Raw URL: `https://libgen.vg/index.php?req=<query>&res=100`
- **Requires a browser User-Agent** (otherwise nginx default page / robot block).
- **Search semantics (verified):** whole-word, case-insensitive. Stems do NOT work reliably
(`происхожден` → 0 hits; full `Происхождение` → many). Multi-word = AND.
→ Use **full inflected words** (nominative, as stored in titles) and 2–3 distinctive words.
- Author search works (`req=Нойманн`) but pulls homonyms — filter by first name in results.
- Pagination: `&page=N` works.
- Edition page: `https://libgen.vg/edition.php?id=<id>`.
### 3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
- URL: `https://cogito-shop.com/search/?q=<query>` (param is **q**, not query)
- Matches author AND title. Good for "is it sellable now in RU".
### 4. SearXNG (local) — cross-verification
- Tool: `tools/sx.sh "query"` (JSON API at http://localhost:8888; the web_search tool blocks localhost)
- Use to verify RU↔EN title pairing in a third party's words, resolve transliterations.
- Keep queries ≤ 3–4 words; «…» quotes only on distinctive RU titles. Unstable relevance —
drop junk queries, don't retry blindly.
### 5. ru.wikipedia / en.wikipedia — author dossiers
- ru-wiki articles often have «…на русском» sections listing actual RU titles.
- en-wiki for the EN bibliography (identity anchor: 3–5 flagship titles).
### 6. livelib.ru author pages — all RU works of a known author
- e.g. `livelib.ru/author/<id>-<slug>` lists every RU edition.
### Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)
- other libgen mirrors (libgen.lol → error page; .vg works with browser UA)
## Identity check protocol (MANDATORY before trusting a hit)
1. **Life dates** in RSL author field (e.g. `Нойманн, Эрих [1905-1960]`).
2. **EN↔RU cross-mapping:** the found person's *other* RU books must map to known EN titles
(e.g. found «Страх феминного» + «Любовь и Душа» → EN *Feminine* + *Love and its Opposites*
→ same Erich Neumann).
3. **Co-publisher signature:** known translation pipelines (Vakler/Рефл-бук 1990s,
Касталия, Прайм-Еврознак, АСТ/ЭКСМО, Питер «Мастера психологии»).
4. **Content spot-check** of the actual target title (open edition page / TOC) when equating
an RU title with an EN title that sounds different
(e.g. *The Archetypal World of Henry Moore* ⇄ «Искусство и творческое бессознательное» — verify it is about Moore).
5. **Homonym rejection:** different first name / field / era → reject or mark ambiguous
(trap found: Верена Каст ≠ Вернер Каст; «Каст» search returns Verena Kast books).
## Workflow (per section)
**Phase 0 — Author dossiers** (`authors/`): EN identity + RU surname candidates.
**Phase 1 — Author sweep:** per author: `rsl.py "<First> <Last>"` → `lg.py "<Last>"` →
`cogito q=<Last>`. Apply identity check. Fill author dossier.
**Phase 2 — Title-gap sweep** for still-unmatched books: `rsl.py "<RU title words>"`
(check the author field!), `lg.py` 2–3 whole-word title combos, cogito.
Negatives: one SearXNG `<EN title> «русский перевод»` before declaring ❌.
**Phase 3 — Edition collection:** ALL editions from ALL sources (title variants! books are
reissued under slightly different names), with full metadata.
**Phase 4 — Book notes + commits.**
Special cases:
- **Jung's CW volumes:** translated and widely available — don't deep-search; note the RU
volume for the item (CW15 = «Дух в человеке, искусстве и литературе», Харвест 2003;
Red Book = «Красная книга (Liber Novus)», Касталия 2025, trans. О. Комков).
- **Essays/chapters** (e.g. Jaffé in *Man and his Symbols*): mark 🔶 if only the parent
book/anthology is available in RU.
## Politeness
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.