fatcat - [no description]

	Commit message (Collapse)	Author	Age	Files	Lines
*	web: fix API URL link for review pages of entities	Bryan Newbold	2021-11-17	1	-2/+2
\|
*	updated notes on possible cleanups	Bryan Newbold	2021-11-17	1	-4/+27
\|
*	ISSN-L dupes check: output all matches	Bryan Newbold	2021-11-17	1	-1/+1
\|
*	document cleanups run this week	Bryan Newbold	2021-11-12	5	-0/+244
\|
*	web: handle ES non-int error codes better	Bryan Newbold	2021-11-12	1	-9/+12
\|
*	Merge branch 'bnewbold-import-refactors' into 'master'	bnewbold	2021-11-11	27	-1599/+874
\|\ \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	import refactors and deprecations Some of these are from old stale branches (the datacite subject metadata patch), but most are from yesterday and today. Sort of a hodge-podge, but the general theme is getting around to deferred cleanups and refactors specific to importer code before making some behavioral changes. The Datacite-specific stuff could use review here. Remove unused/deprecated/dead code: - cdl_dash_dat and wayback_static importers, which were for specific early example entities and have been superseded by other importers - "extid map" sqlite3 feature from several importers, was only used for initial bulk imports (and maybe should not have been used) Refactors: - moved a number of large datastructures out of importer code and into a dedicated static file (`biblio_lookup_tables.py`). Didn't move all, just the ones that were either generic or very large (making it hard to read code) - shuffled around relative imports and some function names ("clean_str" vs. "clean") Some actual behavioral changes: - remove some Datacite-specific license slugs - stop trying to fix double-slashes in DOIs, that was causing more harm than help (some DOIs do actually have double-slashes!) - remove some excess metadata from datacite 'extra' fields
\| *	update datacite tests for license slug changes	Bryan Newbold	2021-11-10	2	-8/+7
\| \| \| \| \| \| \| \| \| \|	Use datacite-specific wrapper function, and remove a couple non-OA/TDM-limited licenses.
\| *	improve lookup_license_slug helper and lookup table	Bryan Newbold	2021-11-10	2	-56/+62
\| \|
\| *	refactor importer metadata tables into separate file; move some helpers around	Bryan Newbold	2021-11-10	10	-702/+682
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	- MAX_ABSTRACT_LENGTH set in a single place (importer common) - merge datacite license slug table in to common table, removing some TDM-specific licenses (which do not apply in the context of preserving the full work)
\| *	importers: refactor imports of clean() and other normalization helpers	Bryan Newbold	2021-11-10	12	-95/+104
\| \|
\| *	remove cdl_dash_dat and wayback_static importers	Bryan Newbold	2021-11-10	4	-596/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Cleaning out dead code. These importers were used to create demonstration fileset and webcapture entities early in development. They have been replaced by the fileset and webcapture ingest importers.
\| *	datacite import: store less subject metadata	Bryan Newbold	2021-11-10	1	-1/+7
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Many of these 'subject' objects have the equivalent of several lines of text, with complex URLs that don't compress well. I think it is fine we have included these thus far instead of parsing more deeply, but going forward I don't think this nested 'extra' metadata is worth the database space.
\| *	add notes about 'double slash in DOI' issue	Bryan Newbold	2021-11-09	1	-0/+46
\| \|
\| *	importers: use clean_doi() in many more (all?) importers	Bryan Newbold	2021-11-09	6	-12/+29
\| \|
\| *	clean_doi: stop mutating double-slash DOIs, except for 10.1037 prefix	Bryan Newbold	2021-11-09	1	-1/+2
\| \|
\| *	remove deprecated extid sqlite3 lookup table feature from importers	Bryan Newbold	2021-11-09	10	-203/+10
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This was used during initial bulk imports, but is no longer used and could create serious metadata problems if used accidentially. In retrospect, it also made metadata provenance less transparent, and may have done more harm than good overall.
* \|	Merge branch 'bnewbold-cleanups-nov2021' into 'master'	bnewbold	2021-11-11	9	-1/+1504
\|\ \ \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Fatcat metadata cleanups/fixups, November 2021 Three cleanups implemented in this branch: - update non-lowercase DOIs on releases (couple hundred thousand entities) - fix incorrectly imported file/release pairs, on the file entity side (~250k entities) - expand truncated wayback URL timestamps in file entities (up to 10 million entities) Instead of proposals, there are documents for each cleanup in `notes/cleanups/`. Have done spot testing of tens of thousands of entities each in QA, and confident about running in production. Plan is to run updates in the order above. DOI and bugfix updates will go fairly fast; the wayback timestamp updates will go slower, and result in large re-indexing load both in fatcat and scholar, because both release and work entities will get triggered for update when file entities are updated.
\| * \|	wayback ts cleanup: one more filter tweak	Bryan Newbold	2021-11-09	1	-1/+2
\| \| \|
\| * \|	update cleanups notes	Bryan Newbold	2021-11-09	2	-0/+72
\| \| \|
\| * \|	file/release bugfix: handle files with multiple edits	Bryan Newbold	2021-11-09	1	-6/+6
\| \| \|
\| * \|	cleanups: add more state=active checks	Bryan Newbold	2021-11-09	2	-0/+8
\| \| \|
\| * \|	update link source filters in file/release bugfix	Bryan Newbold	2021-11-09	1	-2/+8
\| \| \|
\| * \|	initial file/release bugfix cleanup worker and notes	Bryan Newbold	2021-11-09	2	-0/+375
\| \| \|
\| * \|	updates to lowercase DOI cleanup	Bryan Newbold	2021-11-09	2	-7/+86
\| \| \|
\| * \|	lowercase DOI lint and check entity status	Bryan Newbold	2021-11-09	1	-4/+5
\| \| \|
\| * \|	more iteration on short wayback timestamp cleanup	Bryan Newbold	2021-11-09	3	-4/+129
\| \| \|
\| * \|	lint: minor import tweak	Bryan Newbold	2021-11-09	1	-1/+1
\| \| \|
\| * \|	cleanups: tweaks to wayback CDX cleanup scripts	Bryan Newbold	2021-11-09	2	-6/+21
\| \| \|
\| * \|	cleanups: initial lowercase DOI cleanup script	Bryan Newbold	2021-11-09	1	-0/+145
\| \| \|
\| * \|	wayback short ts: another regression test, and some small fmt/tweaks	Bryan Newbold	2021-11-09	1	-3/+38
\| \| \|
\| * \|	wayback cleanup: actually update entity	Bryan Newbold	2021-11-09	1	-2/+4
\| \| \|
\| * \|	imports: generic file cleanup removes exact duplicate URLs	Bryan Newbold	2021-11-09	1	-0/+9
\| \| \|
\| * \|	wayback short ts: add regression test for dupe URLs	Bryan Newbold	2021-11-09	1	-0/+44
\| \| \|
\| * \|	short wayback ts: initial cleanup script implementation	Bryan Newbold	2021-11-09	1	-0/+251
\| \| \|
\| * \|	wayback timestamps: updates to handle 4-digit case	Bryan Newbold	2021-11-09	2	-11/+108
\| \| \|
\| * \|	start work on wayback short-timestamp cleanup	Bryan Newbold	2021-11-09	2	-0/+238
\| \|/
* \|	update crawlability docs	Bryan Newbold	2021-11-10	1	-1/+9
\| \|
* \|	sitemap generation improvements	Bryan Newbold	2021-11-10	2	-1/+2
\| \|
* \|	start notes/proposal about 'crawlability' improvements	Bryan Newbold	2021-11-10	1	-0/+68
\| \|
* \|	pubmed: allow updates if PMCID does not exist yet	Bryan Newbold	2021-11-10	1	-1/+6
\|/ \| \| \| \| \| \| \| \| \| \|	The intent of this change is to start updating Pubmed metadata records when a PMCID has been assigned, but that ext_id hasn't been recorded in fatcat yet. It is likely that this change will result in some additional duplicate PMCIDs in the catalog. But the principle is that the PMID is the primary pubmed identifier, and all records with a PMID should have the PMCID that pubmed indicates, even if there exists another incorrect record.
*	update CHANGELOG for recent development	Bryan Newbold	2021-11-05	1	-0/+26
\|
*	python tests: verify array sort order	Bryan Newbold	2021-11-05	4	-20/+18
\| \| \| \| \| \| \|	In a couple cases (eg, filesets), had made tests agnostic to sort order, because the sort order was not stable. In other cases, simply small cleanups and comment improvements.
*	api: add SQL 'ORDER BY' to many reads to stabilize API array ordering	Bryan Newbold	2021-11-05	1	-3/+14
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	The hope is to make things like file entity URLs, fileset manifests, and other arrays in the JSON API "stable", meaning that if you create an entity with a list of a given order, a read back (in any environment, including prod/QA, bulk dumps, etc) will return the array with the same sort order. This was informally happening most of the time, but occasionally not (!) Assumption is that these sorts will have little or no performance impact, as the common case is less than a dozen elements, and the hard cases are a few thousand at most, and there is already a sorted index.
*	enable type annotation checking with flake8 by default ('make lint')	Bryan Newbold	2021-11-03	1	-4/+2
\|
*	cleanups: create a separate JsonLinePusher for cleanup workers (distinct ↵	Bryan Newbold	2021-11-03	3	-4/+20
\| \| \| \|	base class)
*	facat_import.py: work around corner case in run_cdl_dash_dat()	Bryan Newbold	2021-11-03	1	-1/+1
\|
*	datacite importer: remove unused 'year_only' variable	Bryan Newbold	2021-11-03	1	-2/+3
\|
*	web: work around remaining type annotation issues	Bryan Newbold	2021-11-03	2	-11/+15
\|
*	ignore type errors in cors.py (third party code)	Bryan Newbold	2021-11-03	1	-2/+2
\|
*	web: fix bytes/text warning logging	Bryan Newbold	2021-11-03	1	-3/+3
\| \| \| \|	Minor issue. Caught by type checking