sec01: round 2 closed — 0 new rescues; cache processed catalogs in data/catalogs/

- data/catalogs/: cogito Jungian series (518 titles, 30 pp.), litres 140409 (24 JSON-LD),
  OPP/MISP reading list (55 entries) — source URLs + dates in headers
- author-level check of 16 ❌ against cogito shop search (11 author queries) + full 518-title
  catalog diff: 0 matches; noise rejected (Shamdasani≠Casement, Jung 'Нераскрытая самость'
  ≠ McFarland-Solomon, Berdyaev essay)
- AGENTS.md: catalog caching convention
This commit is contained in:
Dmitry Kokorin 2026-09-18 09:28:33 +03:00
parent 6133d60438
commit 2157d058df
5 changed files with 651 additions and 0 deletions

View file

@ -257,6 +257,13 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
- Anchor titles in result rows contain a raw `>` (title="Хранение<br/>Заказ") — regexes must
not use [^>]+ across the anchor attrs.
## Catalog caching convention
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source
URL + date in the header) so negative verification stays re-runnable without re-crawling.
Examples: `cogito-yungianskaya-series-2026-07-17.md` (30 pp. CRW crawl, 518 titles),
`litres-140409-2026-07-17.md`, `opp-misp-reading-list-2026-07-17.md`.
## Not usable / low value
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
- imaton.com (МААП educational publisher, not translations)