|  | Commit message (Collapse) | Author | Age | Files | Lines | 
|---|
| | |  | 
| | |  | 
| | 
| 
| 
| 
| 
| | Should be backwards compatible with old ingest results.
Fixed a bug with glutton ident detection. | 
| | |  | 
| | |  | 
| | |  | 
| | |  | 
| | |  | 
| | |  | 
| | |  | 
| | |  | 
| |\  
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | Pipfile.lock is broken.
* martin-datacite-import: (68 commits)
  datacite: pass in doi into factored out method
  datacite: reformat test cases and use jq . --sort-keys
  datacite: factor out contributor handling
  datacite: catch type mismatch in language detection
  datacite: adjust tests for release_month
  datacite: name extra.month, extra.release_month
  datacite: mark additional files as stub
  datacite: CCDC are entries, mostly
  datacite: use more specific release_type, if possible
  datacite: ignore certain names
  datacite: over 3% records have the same title: stub
  datacite: fill a few more release_type gaps
  datacite: adding datacite-specific extra metadata
  datacite: apply pylint suggestions
  datacite: fix typos
  datacite: set release_stage to published by default
  datacite: month field should be top-level
  datacite: include month in extra
  datacite: indicate mismatched file in test
  datacite: clean abstracts, use unknown value tokens
  ... | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | | Use values from:
* attributes.creators[]
* attributes.contributors[] | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | The GBIF (https://www.gbif.org/) deposits most records under the titles:
* 599243 GBIF Occurrence Download
* 41176 Occurrence Download
Mark them as "stub" for the moment
(https://guide.fatcat.wiki/entity_release.html#release_type-vocabulary). | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | * citeproc: http://docs.citationstyles.org/en/stable/specification.html#appendix-iii-types
* resourceTypeGeneral: https://schema.datacite.org/meta/kernel-4.0/doc/DataCite-MetadataKernel_v4.0.pdf#page=32
* resourceType: uncontrolled, over 170000 distinct values, frequent:
null, Dataset, JournalArticle, PGRFA Material, Journal Article,
Dataset/UNITE Species Hypothesis, ...
General frequency:
* "attributes.types": 18210075,
* "attributes.types.ris": 18058890,
* "attributes.types.bibtex": 18058888,
* "attributes.types.citeproc": 18058890,
* "attributes.types.schemaOrg": 18058929,
* "attributes.types.resourceType": 12737988,
* "attributes.types.resourceTypeGeneral": 16576139, | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | * attributes.metadataVersion
* attributes.schemaVersion
* attributes.version (source dependent values, follows suggestions in
https://schema.datacite.org/meta/kernel-4.3/doc/DataCite-MetadataKernel_v4.3.pdf#page=26,
but values vary)
Furthermore:
* attributes.types.resourceTypeGeneral
* attributes.types.resourceType | 
| | | |  | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | 
| | | Set to `None` only if there is no publisher yet.
Docs: https://support.datacite.org/docs/doi-states | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | | > include release_month as a top-level extra field [...] to
auto-populate the schema field from that | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | Datacite defines placeholders for unknown values:
* https://support.datacite.org/docs/schema-values-unknown-information-v43
Clean abstracts. | 
| | | |  | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | 
| | | > always include extra values for the respective DOI registrars
(datacite, crossref, jalc), even if they are empty ({}), to be used as a
flag so we know which DOI registrar supplied the metadata. | 
| | | |  | 
| | | 
| | 
| | 
| | | As [...] we will soon add support for release_month field in the release schema. | 
| | | |  | 
| | | |  | 
| | | 
| | 
| | 
| | | Estimated time for a single call is in the order of 50ms. | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | 
| | 
| | | > The convention for display_name and raw_name is to be how the name
would normally be printed, not in index form (surname comma given_name).
So we might need to un-encode names like "Tricart, Pierre".
Use an additional `index_form_to_display_name` function to convert index
from to display form, heuristically. | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | |  | 
| | | 
| | 
| | 
| | 
| | 
| | | Example of a non-ascii doi:
* https://doi.org/10.13125/américacrítica/3017 | 
| | | 
| | 
| | 
| | 
| | 
| | 
| | | address issue with EN DASH DOI.
> "external identifier doesn't match required pattern for a DOI (expected,
eg, '10.1234/aksjdfh'): 10.25513/1812-3996.2017.1.34–42" | 
| | | |  | 
| | | |  | 
| | | |  |