jung/AGENTS.md

132 lines
7.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Jung Reading List → Russian Editions: Research Notes
## Goal
For every book in the ISAP Zurich reading list (https://isapzurich.com/en/library/reading-list),
find the **Russian edition(s)**: correct Russian title, publisher, year, translator, ISBN,
libgen download page (if exists), and РГБ (RSL) bibliographic record (if exists).
We work **section by section**. Current: **Section 04 — Pictures** (`sections/04-pictures/`).
Sections 1–3, 5–7 are skeletons for later.
## Repository layout
```
AGENTS.md # this file — plan, conventions, source status
CONTEXT.md # misc findings that fit no other category
tools/ # search helpers (lg.py, rsl.py, sx.sh)
authors/ # one dossier per author (created on demand)
sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged into the book file)
```
## Conventions
- **One file per book**, named `NN-<slug>.md` (NN = order in the reading list).
- **Status marker** at top of each book file:
- `⬜ not searched yet`
- `🔎 searching in progress`
- `✅ RU edition(s) found`
- `🔶 partial` (e.g. only an essay/chapter available in RU)
- `❌ no RU edition found (verified)`
- **Book file contents:** human-readable editions table + raw **RSL JSON** block +
libgen edition URLs (`https://libgen.vg/edition.php?id=...` — contains download links;
direct file links rotate, don't store them) + notes (identity check, RU↔EN title pairing source).
- **Author dossiers** (`authors/<lastname>-<firstname>.md`) hold: EN identity
(3–5 flagship EN titles, dates), RU surname candidates (all plausible transliterations,
German names often have several: Нойманн/Нойман/Нейманн), verified RU works with EN mapping,
homonym warnings.
- **Commit per completed book** (or when meaningful author info is gained), message like:
`sec04: <book> — <finding>`.
- Update `CONTEXT.md` for anything useful that fits no category.
## Sources and their status (verified 2026)
### 1. РГБ (Russian State Library) via libgen.vg biblioservice — PRIMARY for existence + editions
- Tool: `tools/rsl.py "Full Name"` (JSON, polite, 3s between pages)
- Raw URL: `https://libgen.vg/biblioservice.php?value=<query>&type=rsl&format=json`
- Rich records: title parts, **author with life dates** `[1905-1960]` (great identity anchor),
translator, publisher, city, year, pages, ISBN, series, UDC tags, even TOC field.
- **Limitations:** no pagination (fixed first page) → always query by **full first+last name**
for people, or 2–3 distinctive words. Loose matching (searches author AND title fields).
- Other `type=` values on this mirror are **dead**: worldcat, googlebooks, isbndb, udc
(all return "Nothing found" even for control queries). `type=isbn` only decodes ISBN structure.
### 2. libgen.vg full-text search — PRIMARY for readable editions
- Tool: `tools/lg.py "query"` (whole words only, res=100, prints pagination hint)
- Raw URL: `https://libgen.vg/index.php?req=<query>&res=100`
- **Requires a browser User-Agent** (otherwise nginx default page / robot block).
- **Search semantics (verified):** whole-word, case-insensitive. Stems do NOT work reliably
(`происхожден` → 0 hits; full `Происхождение` → many). Multi-word = AND.
→ Use **full inflected words** (nominative, as stored in titles) and 2–3 distinctive words.
- Author search works (`req=Нойманн`) but pulls homonyms — filter by first name in results.
- Pagination: `&page=N` works.
- Edition page: `https://libgen.vg/edition.php?id=<id>`.
### 3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
- URL: `https://cogito-shop.com/search/?q=<query>` (param is **q**, not query)
- Matches author AND title. Good for "is it sellable now in RU".
### 4. SearXNG (local) — cross-verification
- Tool: `tools/sx.sh "query"` (JSON API at http://localhost:8888; the web_search tool blocks localhost)
- Use to verify RU↔EN title pairing in a third party's words, resolve transliterations.
- Keep queries ≤ 3–4 words; «…» quotes only on distinctive RU titles. Unstable relevance —
drop junk queries, don't retry blindly.
### 5. ru.wikipedia / en.wikipedia — author dossiers
- ru-wiki articles often have «…на русском» sections listing actual RU titles.
- en-wiki for the EN bibliography (identity anchor: 3–5 flagship titles).
### 6. livelib.ru author pages — all RU works of a known author
- e.g. `livelib.ru/author/<id>-<slug>` lists every RU edition (server-rendered, scrapable).
- Author ID: find via SearXNG `site:livelib.ru <author>`. The /search page is JS-only (nope).
### 7. cogito-shop person pages — shop stock per author
- `cogito-shop.com/person/<first>_<last>/` (transliterated, first-name first,
e.g. `/person/mariya_luiza_fon_frants/`, `/person/yaffe_aniela/`).
- Slug discoverable from the shop's search page (results contain person links) or guessable.
- Lists every edition the shop carries (title variants!). Product pages lack full
metadata in HTML (JS-rendered) — use RSL/livelib for publisher/year/ISBN.
### Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)
- bookmate.com (403 bot protection), gtmarket.ru (JS app, no HTML content)
- other libgen mirrors (libgen.lol → error page; .vg works with browser UA)
## Identity check protocol (MANDATORY before trusting a hit)
1. **Life dates** in RSL author field (e.g. `Нойманн, Эрих [1905-1960]`).
2. **EN↔RU cross-mapping:** the found person's *other* RU books must map to known EN titles
(e.g. found «Страх феминного» + «Любовь и Душа» → EN *Feminine* + *Love and its Opposites*
→ same Erich Neumann).
3. **Co-publisher signature:** known translation pipelines (Vakler/Рефл-бук 1990s,
Касталия, Прайм-Еврознак, АСТ/ЭКСМО, Питер «Мастера психологии»).
4. **Content spot-check** of the actual target title (open edition page / TOC) when equating
an RU title with an EN title that sounds different
(e.g. *The Archetypal World of Henry Moore* ⇄ «Искусство и творческое бессознательное» — verify it is about Moore).
5. **Homonym rejection:** different first name / field / era → reject or mark ambiguous
(trap found: Верена Каст ≠ Вернер Каст; «Каст» search returns Verena Kast books).
## Workflow (per section)
**Phase 0 — Author dossiers** (`authors/`): EN identity + RU surname candidates.
**Phase 1 — Author sweep:** per author: `rsl.py "<First> <Last>"` → `lg.py "<Last>"` →
`cogito q=<Last>`. Apply identity check. Fill author dossier.
**Phase 2 — Title-gap sweep** for still-unmatched books: `rsl.py "<RU title words>"`
(check the author field!), `lg.py` 2–3 whole-word title combos, cogito.
Negatives: one SearXNG `<EN title> «русский перевод»` before declaring ❌.
**Phase 3 — Edition collection:** ALL editions from ALL sources (title variants! books are
reissued under slightly different names), with full metadata.
**Phase 4 — Book notes + commits.**
Special cases:
- **Jung's CW volumes:** translated and widely available — don't deep-search; note the RU
volume for the item (CW15 = «Дух в человеке, искусстве и литературе», Харвест 2003;
Red Book = «Красная книга (Liber Novus)», Касталия 2025, trans. О. Комков).
- **Essays/chapters** (e.g. Jaffé in *Man and his Symbols*): mark 🔶 if only the parent
book/anthology is available in RU.
## Politeness
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.