pipeline v3: RU-title candidates (rutitles+ru-dict+pymorphy3), batch sweep driver, OpenAlex/Crossref

- tools/rutitles.py: EN->RU token translation (data/ru-dict.tsv, 540+ pairs, Kogito/Castalia
  conventions) + Jaccard diff against data/ru-titles.jsonl (692 titles: Kogito 518, Litres 24,
  OPP 55, flibusta a/5272+57639+118921+122883+193690+193689). Matching is lemmatized
  (pymorphy3) — case endings handled: 'великой матери' -> 'Великая мать' 1.25.
  'The Great Mother' -> 'Великая мать' 1.25 top hit; 'The Symbolic Quest' -> correct negative.
- tools/sweep.py: batch driver for Phases 1-2 (rsl+lg+alib+flib+cogito+SearXNG-OZON-snippet
  parse+rutitles diff per item; --nlr optional). Query log data/queries.log (JSONL).
  Fixed: cogito nav-menu leak (parse bx_product_item only), libgen robot-block (lg.py curl fallback).
- tools/oa.py: OpenAlex + Crossref (Phase 0 identity/ISBN, no key; found Margaret Wilkinson,
  Karen Evers-Fahey, Symbolic Quest Princeton ISBN).
- data/sweeps/01-fundamentals/input.tsv: 16 ❌ items loaded; background sweep running.
- AGENTS.md: source 10b (oa.py), flibusta .su = reduced mirror (dropped from pipeline),
  pipeline v3 section (rutitles/oa/sweep/pymorphy3 note: pymorphy2 broken on py3.12).
- data/ru-titles.jsonl committed as the RU market universe asset.
This commit is contained in:
Dmitry Kokorin 2026-09-18 10:56:36 +03:00
parent 2157d058df
commit 5441dddbb8
11 changed files with 2101 additions and 3 deletions

View file

@ -121,13 +121,22 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
- Author HTML: `/a/<id>` (e.g. a/57639 = фон Франц, a/118921 = Нойманн); book page `/b/<id>`
- **Direct downloads:** `/b/<id>/fb2|epub|mobi|html|txt` (pdf/doc as «/b/<id>/download» links on the pages)
- **Quirks (2026-07-17):** `fb2+zip` format actually returns a ZIP (unpack: 1 fb2 + cover inside); `/b/<id>/pdf` sometimes 302-redirects to the book HTML page (a ~20 KB stub) when the file isn't served directly — check downloaded «pdf»s with `file`, fall back to libgen (same editions usually there) or the fb2 twin.
- .su search: `https://flibusta.su/booksearch/?ask=<query>` (param is **ask**); HTML links
`/book/<id>-<slug>`, `/author/<id>-<slug>`. API endpoints blocked on .su (api/search.php → 403).
- .su = **mirror of .is with reduced functionality** (user, 2026-07-17) — use .is OPDS only.
.su `booksearch/?ask=` exists but adds nothing; API endpoints blocked (api/search.php → 403).
- Value: (1) RU full-text downloads for the download phase; (2) author-book-list = cheap author sweep;
(3) negative verification (all same-name authors listed). Verified 2026-07-09 (.su): no «Выявляющий
образ», no «Сэндплей» titles, no «Тест Дерево», no «Искусство и творческое бессознательное» (as titled).
Flibusta search is fuzzy (whole-phrase not required).
### 10b. OpenAlex + Crossref (`tools/oa.py`) — Phase 0 identity/ISBN (free, no key, better than OL for modern academic books)
- `oa.py "Exact EN Title"` (OpenAlex works: full author names, year, DOI, publisher),
`oa.py --crossref "Title words"` (Crossref: **ISBN list** + publisher + year),
`oa.py --author "Last, First"` (OpenAlex author entity), `oa.py both "Title"`.
- Verified 2026-07-17: Evers-Fahey full name "Karen Evers-Fahey" + Routledge 2016;
"Coming into Mind" = **Margaret** Wilkinson (not Michael!) 9781317710578;
"The Symbolic Quest" Whitmont 9780691213187 (Princeton).
- Use in Phase 0 INSTEAD of / alongside OL; OL still primary for pre-1990 titles.
### 11. Open Library — the "English RSL" (author works lists, full names, ISBNs)
- Tool: `tools/ol.py "author:Lastname, First"` or `tools/ol.py 'title:"Exact EN Title"'`
- `openlibrary.org/search.json?q=...&fields=title,author_name,publish_year,language,publisher,isbn`
@ -264,6 +273,27 @@ URL + date in the header) so negative verification stays re-runnable without re-
Examples: `cogito-yungianskaya-series-2026-07-17.md` (30 pp. CRW crawl, 518 titles),
`litres-140409-2026-07-17.md`, `opp-misp-reading-list-2026-07-17.md`.
## Pipeline v3 (2026-07-17) — RU-title candidates & batch sweeps
- **`tools/rutitles.py`** — the "diff, not search" tool:
- `build` — rebuild `data/ru-titles.jsonl` from `data/catalogs/*` + preserve flibusta entries
- `addflib <authorId>` — append flibusta.is authorall titles to the DB
- `translate "EN Title"` — RU token bag via `data/ru-dict.tsv` (540+ EN→RU pairs, phrases first,
observed Kogito/Castalia/Peter conventions) + list of UNTRANSLATED words (finish manually)
- `match "RU tokens" [N]` / `both "EN Title"` — Jaccard over **pymorphy3-lemmatized** tokens
(case endings handled: «великой матери» → «Великая мать» 1.25; order-agnostic)
- DB (2026-07-17): 692 titles (Когито-518 + ЛитРес-24 + ОПП/МИСП-55 + flib a/5272, 57639, 118921,
122883, 193690, 193689)
- **Limits (honest):** generic RU titles (Wolff «Введение в основы…») invisible — publisher
sweep (Phase 2b) stays a separate phase. Translation ≠ identity: verify real hits.
- **pymorphy3** installed (user-site): lemmatization for RU token matching.
⚠ pymorphy2 is BROKEN on Python 3.12 (`inspect.getargspec` removed) — use pymorphy3 only.
- **`tools/oa.py`** — OpenAlex/Crossref (Phase 0, see source 10b).
- **flibusta .su dropped** from the search pipeline (mirror of .is, reduced functionality).
- **`tools/sweep.py`** — batch driver for Phases 1–2 (per author/item: rsl + lg + alib + flib +
cogito + SearXNG-OZON-snippet parse + rutitles diff; optional `--nlr` for НРБ via CRW, slow).
Query log: `data/queries.log` (JSONL). Sweep output: `data/sweeps/secNN/<nn>-<slug>.md`.
## Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)