pipeline v3: RU-title candidates (rutitles+ru-dict+pymorphy3), batch sweep driver, OpenAlex/Crossref

- tools/rutitles.py: EN->RU token translation (data/ru-dict.tsv, 540+ pairs, Kogito/Castalia
  conventions) + Jaccard diff against data/ru-titles.jsonl (692 titles: Kogito 518, Litres 24,
  OPP 55, flibusta a/5272+57639+118921+122883+193690+193689). Matching is lemmatized
  (pymorphy3) — case endings handled: 'великой матери' -> 'Великая мать' 1.25.
  'The Great Mother' -> 'Великая мать' 1.25 top hit; 'The Symbolic Quest' -> correct negative.
- tools/sweep.py: batch driver for Phases 1-2 (rsl+lg+alib+flib+cogito+SearXNG-OZON-snippet
  parse+rutitles diff per item; --nlr optional). Query log data/queries.log (JSONL).
  Fixed: cogito nav-menu leak (parse bx_product_item only), libgen robot-block (lg.py curl fallback).
- tools/oa.py: OpenAlex + Crossref (Phase 0 identity/ISBN, no key; found Margaret Wilkinson,
  Karen Evers-Fahey, Symbolic Quest Princeton ISBN).
- data/sweeps/01-fundamentals/input.tsv: 16 ❌ items loaded; background sweep running.
- AGENTS.md: source 10b (oa.py), flibusta .su = reduced mirror (dropped from pipeline),
  pipeline v3 section (rutitles/oa/sweep/pymorphy3 note: pymorphy2 broken on py3.12).
- data/ru-titles.jsonl committed as the RU market universe asset.
This commit is contained in:
Dmitry Kokorin 2026-09-18 10:56:36 +03:00
parent e7b35223af
commit ba2a10790c
11 changed files with 2101 additions and 3 deletions

View file

@ -10,10 +10,24 @@ def fetch(url, tries=3):
for i in range(tries):
try:
req = urllib.request.Request(url, headers={"User-Agent": UA, "Accept-Language": "ru-RU,ru;q=0.9"})
return urllib.request.urlopen(req, timeout=90).read().decode("utf-8", "replace")
html = urllib.request.urlopen(req, timeout=90).read().decode("utf-8", "replace")
# robot-block detection: nginx default page has no req= result table
if 'libgen' not in html.lower()[:3000] and 'Search Result' not in html and 'result' not in html.lower()[:3000]:
time.sleep(15)
continue
return html
except Exception as e:
last = e
time.sleep(5 + 5 * i)
# curl fallback (works when urllib is fingerprinted/blocked)
import subprocess
try:
r = subprocess.run(["curl", "-s", "-A", UA, "--max-time", "90", url],
capture_output=True, text=True, timeout=120)
if r.stdout:
return r.stdout
except Exception:
pass
raise SystemExit(f"fetch failed after {tries} tries: {last}")
def main():