fatcat - [no description]

	Commit message (Collapse)	Author	Age	Files	Lines
*	chocula importer: handle not-upper-case ISSNs	Bryan Newbold	2021-11-30	1	-2/+6
\|
*	chocula importer: handle broken ISSNs in extra metadata	Bryan Newbold	2021-11-30	1	-2/+7
\|
*	chocula importer: tweak counting, conditions for doing updates	Bryan Newbold	2021-11-30	1	-15/+7
\|
*	chocula importer: move issne/issnp 'extra' to top-level fields if doing updates	Bryan Newbold	2021-11-30	1	-0/+6
\|
*	chocula: don't do name cleanups in importer	Bryan Newbold	2021-11-30	1	-8/+2
\| \| \| \|	This kind of cleanup should be done in 'chocula' instead.
*	codespell fixes in python code (comments)	Bryan Newbold	2021-11-24	1	-2/+2
\|
*	Merge branch 'bnewbold-import-refactors' into 'master'	bnewbold	2021-11-11	16	-1380/+146
\|\ \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	import refactors and deprecations Some of these are from old stale branches (the datacite subject metadata patch), but most are from yesterday and today. Sort of a hodge-podge, but the general theme is getting around to deferred cleanups and refactors specific to importer code before making some behavioral changes. The Datacite-specific stuff could use review here. Remove unused/deprecated/dead code: - cdl_dash_dat and wayback_static importers, which were for specific early example entities and have been superseded by other importers - "extid map" sqlite3 feature from several importers, was only used for initial bulk imports (and maybe should not have been used) Refactors: - moved a number of large datastructures out of importer code and into a dedicated static file (`biblio_lookup_tables.py`). Didn't move all, just the ones that were either generic or very large (making it hard to read code) - shuffled around relative imports and some function names ("clean_str" vs. "clean") Some actual behavioral changes: - remove some Datacite-specific license slugs - stop trying to fix double-slashes in DOIs, that was causing more harm than help (some DOIs do actually have double-slashes!) - remove some excess metadata from datacite 'extra' fields
\| *	refactor importer metadata tables into separate file; move some helpers around	Bryan Newbold	2021-11-10	8	-621/+25
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	- MAX_ABSTRACT_LENGTH set in a single place (importer common) - merge datacite license slug table in to common table, removing some TDM-specific licenses (which do not apply in the context of preserving the full work)
\| *	importers: refactor imports of clean() and other normalization helpers	Bryan Newbold	2021-11-10	12	-95/+104
\| \|
\| *	remove cdl_dash_dat and wayback_static importers	Bryan Newbold	2021-11-10	3	-510/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Cleaning out dead code. These importers were used to create demonstration fileset and webcapture entities early in development. They have been replaced by the fileset and webcapture ingest importers.
\| *	datacite import: store less subject metadata	Bryan Newbold	2021-11-10	1	-1/+7
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Many of these 'subject' objects have the equivalent of several lines of text, with complex URLs that don't compress well. I think it is fine we have included these thus far instead of parsing more deeply, but going forward I don't think this nested 'extra' metadata is worth the database space.
\| *	importers: use clean_doi() in many more (all?) importers	Bryan Newbold	2021-11-09	6	-12/+29
\| \|
\| *	remove deprecated extid sqlite3 lookup table feature from importers	Bryan Newbold	2021-11-09	3	-160/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This was used during initial bulk imports, but is no longer used and could create serious metadata problems if used accidentially. In retrospect, it also made metadata provenance less transparent, and may have done more harm than good overall.
* \|	Merge branch 'bnewbold-cleanups-nov2021' into 'master'	bnewbold	2021-11-11	1	-0/+9
\|\ \ \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Fatcat metadata cleanups/fixups, November 2021 Three cleanups implemented in this branch: - update non-lowercase DOIs on releases (couple hundred thousand entities) - fix incorrectly imported file/release pairs, on the file entity side (~250k entities) - expand truncated wayback URL timestamps in file entities (up to 10 million entities) Instead of proposals, there are documents for each cleanup in `notes/cleanups/`. Have done spot testing of tens of thousands of entities each in QA, and confident about running in production. Plan is to run updates in the order above. DOI and bugfix updates will go fairly fast; the wayback timestamp updates will go slower, and result in large re-indexing load both in fatcat and scholar, because both release and work entities will get triggered for update when file entities are updated.
\| * \|	imports: generic file cleanup removes exact duplicate URLs	Bryan Newbold	2021-11-09	1	-0/+9
\| \|/
* /	pubmed: allow updates if PMCID does not exist yet	Bryan Newbold	2021-11-10	1	-1/+6
\|/ \| \| \| \| \| \| \| \| \| \|	The intent of this change is to start updating Pubmed metadata records when a PMCID has been assigned, but that ext_id hasn't been recorded in fatcat yet. It is likely that this change will result in some additional duplicate PMCIDs in the catalog. But the principle is that the PMID is the primary pubmed identifier, and all records with a PMID should have the PMCID that pubmed indicates, even if there exists another incorrect record.
*	datacite importer: remove unused 'year_only' variable	Bryan Newbold	2021-11-03	1	-2/+3
\|
*	datacite: add comment about potential date parsing bug	Bryan Newbold	2021-11-03	1	-0/+1
\|
*	datacite importer: dateparser.date.DateDataParser()	Bryan Newbold	2021-11-03	1	-1/+1
\| \| \| \|	Perhaps this was a change when upgrading 'dateparser'?
*	more involved type wrangling and fixes for importers	Bryan Newbold	2021-11-03	3	-12/+14
\|
*	typing: relatively simple type check fixes	Bryan Newbold	2021-11-03	14	-87/+82
\| \| \| \| \| \| \|	These mostly add new variable names so that existing variables aren't overwritten with a new type; delay coercing '{}' or '[]' to 'None' until the last minute; adding is-not-None checks to conditional clauses; and similar small changes.
*	typing: initial annotations on importers	Bryan Newbold	2021-11-03	22	-274/+443
\| \| \| \| \|	This commit just adds the type annotations, doesn't do fixes to code to make type checking pass.
*	importers: remove unused __main__ routine	Bryan Newbold	2021-11-03	4	-19/+0
\| \| \| \| \| \|	These perhaps were used in initial develoment or testing? fatcat_import.py is the correct way to do these imports, even for testing/development.
*	lint: resolve existing mypy type errors	Bryan Newbold	2021-11-02	3	-22/+27
\| \| \| \| \| \| \| \| \|	Adds annotations and re-workes dataflow to satisfy existing mypy issues, without adding any additional type annotations to, eg, function signatures. There will probably be many more type errors when annotations are all added.
*	re-fix some lint issues after big 'fmt'	Bryan Newbold	2021-11-02	1	-2/+2
\|
*	fmt (black): fatcat_tools/	Bryan Newbold	2021-11-02	22	-2115/+2578
\|
*	python: isort everything	Bryan Newbold	2021-11-02	17	-41/+70
\|
*	arabesque import 'hit' field is 1/0, not true/false	Bryan Newbold	2021-11-02	1	-2/+2
\|
*	lint: simple, safe inline lint fixes	Bryan Newbold	2021-11-02	12	-22/+21
\| \| \| \|	'==' vs 'is'; 'not a in b' vs 'a not in b'; etc
*	lint/fmt: remove all 'import *'	Bryan Newbold	2021-11-02	5	-21/+41
\|
*	re-fmt all the fatcat_tools __init__ files for readability	Bryan Newbold	2021-11-02	1	-17/+39
\|
*	small python tweaks for annotations, imports	Bryan Newbold	2021-11-02	2	-2/+6
\|
*	try some type annotations	Bryan Newbold	2021-11-02	2	-55/+63
\|
*	fix missing variable in fileset ingest	Bryan Newbold	2021-11-02	1	-2/+1
\|
*	WIP: more fileset ingest	Bryan Newbold	2021-10-18	1	-13/+21
\|
*	WIP: rel fixes	Bryan Newbold	2021-10-14	1	-6/+6
\|
*	fileset ingest small tweaks	Bryan Newbold	2021-10-14	1	-21/+36
\|
*	initial implementation of fileset ingest importers	Bryan Newbold	2021-10-14	2	-3/+224
\|
*	generic fileset importer class, with test coverage	Bryan Newbold	2021-10-14	3	-0/+88
\|
*	dblp import: basic support for handles as identifiers	Bryan Newbold	2021-10-13	1	-1/+5
\|
*	dblp import: fix typos in identifier parsing	Bryan Newbold	2021-10-13	1	-2/+1
\|
*	python: partial importer utilization of new schema changes	Bryan Newbold	2021-10-13	3	-6/+18
\|
*	Merge branch 'bnewbold-ingest-tweaks' into 'master'	bnewbold	2021-10-02	3	-39/+106
\|\ \| \| \| \| \| \| \| \|	ingest importer behavior tweaks See merge request webgroup/fatcat!120
\| *	kafka import: optional 'force-flush' mode for some importers	Bryan Newbold	2021-10-01	1	-0/+13
\| \| \| \| \| \| \| \|	Behavior and motivation described in the kafka json import comment.
\| *	new SPN web (html) importer	Bryan Newbold	2021-10-01	2	-27/+81
\| \|
\| *	ingest importer behavior tweaks	Bryan Newbold	2021-10-01	1	-8/+8
\| \| \| \| \| \| \| \| \| \|	- change order of 'want()' checks, so that result counts are clearer - don't require GROBID success for file imports with SPN
\| *	importer common: more verbose logging (with counts)	Bryan Newbold	2021-10-01	1	-4/+4
\| \|
* \|	datacite: skip empty abstracts	Martin Czygan	2021-10-01	1	-1/+4
\|/ \| \| \| \|	Do not add abstracts where `clean` results in the empty string - this violates a constraint: `either abstract_sha1 or content is required`
*	more consistent and defensive lower-casing of DOIs	Bryan Newbold	2021-06-23	2	-1/+6
\| \| \| \| \| \| \|	After noticing more upper/lower ambiguity in production. In particular, we have some old ingest requests in sandcrawler DB, which get re-submitted/re-tried, which have capitalized DOIs in the link source id field.
*	datacite: more careful title string access; fixes sentry #88350	Martin Czygan	2021-06-11	1	-1/+1
\| \| \| \| \|	Caused by a partial "title entry without title" coming first (e.g. just holding, e.g. a language, like: {'lang': 'da'}