aboutsummaryrefslogtreecommitdiffstats
Commit message (Expand)AuthorAgeFilesLines
* update fatcat_file SQL table schema, and add backfill notesBryan Newbold2021-12-071-1/+3
* update fatcat_file SQL table schema, and add backfill notesBryan Newbold2021-12-011-0/+13
* commit old patch crawl notesBryan Newbold2021-12-011-0/+488
* Revert "pipenv: update deps"Bryan Newbold2021-12-012-574/+382
* pipenv: update depsBryan Newbold2021-12-012-382/+574
* add CDX sha1hex lookup/fetch helper scriptBryan Newbold2021-11-301-0/+170
* sandcrawler SQL statsBryan Newbold2021-11-272-12/+425
* codespell typos in README and original RFCBryan Newbold2021-11-242-2/+2
* codespell typos in python (comments)Bryan Newbold2021-11-245-5/+5
* html_meta: actual typo in code (CSS selector) caught by codespellBryan Newbold2021-11-241-1/+1
* codespell fixes in proposalsBryan Newbold2021-11-248-16/+16
* ingest tool: new backfill modeBryan Newbold2021-11-161-0/+76
* make fmtBryan Newbold2021-11-161-1/+1
* SPNv2: make 'resources' optionalBryan Newbold2021-11-161-1/+1
* grobid: handle XML parsing errors, and have them recorded in sandcrawler-dbBryan Newbold2021-11-121-1/+5
* ingest_file: more efficient GROBID metadata copyBryan Newbold2021-11-121-3/+3
* wrap up crossref refs backfill notesBryan Newbold2021-11-101-0/+47
* grobid_tool: helper to process a single fileBryan Newbold2021-11-101-0/+15
* ingest: start re-processing GROBID with newer versionBryan Newbold2021-11-101-2/+6
* simple persist worker/tool to backfill grobid_refsBryan Newbold2021-11-102-0/+62
* grobid: extract more metadata in document TEI-XMLBryan Newbold2021-11-101-0/+5
* grobid: update 'TODO' comment based on reviewBryan Newbold2021-11-041-3/+0
* update crossref/grobid refs generation notesBryan Newbold2021-11-041-4/+96
* crossref grobid refs: another error case (ReadTimeout)Bryan Newbold2021-11-042-5/+11
* db (postgrest): actually use an HTTP sessionBryan Newbold2021-11-041-12/+24
* grobid: use requests sessionBryan Newbold2021-11-041-3/+4
* grobid crossref refs: try to handle HTTP 5xx and XML parse errorsBryan Newbold2021-11-042-5/+33
* grobid: handle weird whitespace unstructured from crossrefBryan Newbold2021-11-041-1/+10
* crossref persist: batch size depends on whether parsing refsBryan Newbold2021-11-042-2/+8
* sql: grobid_refs table JSON as 'JSON' not 'JSONB'Bryan Newbold2021-11-042-3/+3
* grobid refs backfill progressBryan Newbold2021-11-041-1/+43
* record SQL table sizes at start of crossref re-ingestBryan Newbold2021-11-041-0/+19
* start notes on crossref refs backfillBryan Newbold2021-11-041-0/+54
* crossref persist: make GROBID ref parsing an option (not default)Bryan Newbold2021-11-043-9/+33
* add grobid_refs and crossref_with_refs to sandcrawler-db SQL schemaBryan Newbold2021-11-041-0/+21
* glue, utils, and worker code for crossref and grobid_refsBryan Newbold2021-11-044-5/+212
* update grobid refs proposalBryan Newbold2021-11-041-10/+72
* iterated GROBID citation cleaning and processingBryan Newbold2021-11-041-27/+45
* grobid citations: first pass at cleaning unstructuredBryan Newbold2021-11-041-2/+34
* initial proposal for GROBID refs table and pipelineBryan Newbold2021-11-041-0/+63
* initial crossref-refs via GROBID helper routineBryan Newbold2021-11-047-6/+839
* pipenv: bump grobid_tei_xml version to 0.1.2Bryan Newbold2021-11-042-11/+11
* pdftrio client: use HTTP session for POSTsBryan Newbold2021-11-031-1/+1
* workers: use HTTP session for archive.org fetchesBryan Newbold2021-11-031-3/+3
* IA (wayback): actually use an HTTP session for replay fetchesBryan Newbold2021-11-031-2/+3
* SPN reingest: 6 hour minimum, 6 month maxBryan Newbold2021-11-031-2/+2
* sql: fix typo in quarterly (not weekly) scriptBryan Newbold2021-11-031-1/+1
* sql: fixes to ingest_fileset_platform schema (from table creation)Bryan Newbold2021-11-012-12/+12
* updates/corrections to old small.json GROBID metadata example fileBryan Newbold2021-10-271-6/+1
* remove grobid2json helper file, replace with grobid_tei_xmlBryan Newbold2021-10-277-224/+22