jung/tools/rulabel.py
Dmitry Kokorin 7b24100450 sec06: 28 карточек (11✅/2🔶/15❌) + rulabel.py (ru labels как канал канонической кириллицы)
- Кросс-рефы: 01→sec01#05, 02→sec02#06, 03→sec01#32, 05→sec03#09, 06→sec01#15, 11→sec03#20
- Новые ✅: Eliade HRI («История веры и религиозных идей» Акад.проект 2008-09),
  Otto «Священное» СПбГУ 2008, Kerényi «Дионис» ВРС 2007, MacCulloch Э/Эксмо 2018,
  Schimmel «Мир исламского мистицизма» Садра 2012, Scholem 4 изд., von Franz CW6
  («Видения Николая из Флюэ…» Акад.проект 2025, rescue через rutitles diff)
- tools/rulabel.py: Wikidata ru label + aliases = каноническая кириллица
  (VIAF/ISNI/GND/BnF с этой машины недоступны: 403/mёртвые API). Ловушки:
  Циммер не Зиммер, Корбен не Корбин, Шолем не Шулем.
- AGENTS.md: источник 12 дополнен (rulabel.py + статус authority-APIs)
2026-09-24 11:26:20 +03:00

63 lines
2.7 KiB
Python
Executable file
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

#!/usr/bin/env python3
"""tools/rulabel.py — каноническое кириллическое имя автора по Wikidata (ru label + ru aliases).
Идея (2026-09-24, предложил пользователь): authority-записи (VIAF/ISNI/GND) хранят имя во
всех транскрипциях — но с этой машины VIAF (403 + JS-app), ISNI (403), GND (API mёртв,
services.dnb.de → 404), BnF SRU (403) НЕДОСТУПНЫ. Wikidata-ru label/aliases доступны свободно
и для известных людей = тот же канонический вариант (обычно взят из русскоязычных sources).
Использование:
rulabel.py Q44847 # один QID
rulabel.py Q1|Q2|Q3 # батч
rulabel.py --find "Scholem, Gershom" # поиск (ОСТОРОЖНО: top-5, омонимы!)
Выход: <QID>\t<EN label>\t<ru label>\t<ru aliases (| разделены)>
"""
import json, sys, urllib.parse, urllib.request, time
UA = "jung-research/1.0 (research script; contact dmitry@kokorin.org)"
def api(params):
url = "https://www.wikidata.org/w/api.php?" + urllib.parse.urlencode(params)
req = urllib.request.Request(url, headers={"User-Agent": UA})
for attempt in range(4):
try:
return json.load(urllib.request.urlopen(req, timeout=45))
except Exception as e:
if attempt == 3:
raise
time.sleep(15 * (attempt + 1))
def show(qid, e):
lab_en = e.get("labels", {}).get("en", {}).get("value", "")
lab_ru = e.get("labels", {}).get("ru", {}).get("value")
al_ru = [a["value"] for a in e.get("aliases", {}).get("ru", [])]
print(f"{qid}\t{lab_en}\t{lab_ru or '—'}\t{' | '.join(al_ru) or '—'}")
def main():
args = sys.argv[1:]
if not args:
sys.exit(__doc__)
if args[0] == "--find":
query = args[1]
d = api({"action": "wbsearchentities", "search": query, "language": "en",
"type": "item", "limit": 5, "format": "json"})
for r in d.get("search", []):
print(f" {r['id']}\t{r.get('label','?')}\t{r.get('description','')[:70]}")
qids = [r["id"] for r in d.get("search", [])]
else:
qids = args[0].split("|")
if not qids:
return
d = api({"action": "wbgetentities", "ids": "|".join(qids),
"props": "labels|aliases", "languages": "en|ru", "format": "json"})
for qid in qids:
e = d.get("entities", {}).get(qid)
if e:
show(qid, e)
else:
print(f"{qid}\tMISSING\t\t")
time.sleep(2)
if __name__ == "__main__":
main()