diff options
author | Bryan Newbold <bnewbold@archive.org> | 2018-05-08 06:21:29 +0000 |
---|---|---|
committer | Bryan Newbold <bnewbold@archive.org> | 2018-05-08 06:21:29 +0000 |
commit | e566ee1b4e134bfc06284cf77d8d1370df30d53f (patch) | |
tree | f3969054cc5f93608b5c72d41541ea381ef89a6b /pig/tests/files/papers_url_doi.cdx | |
parent | 0c398392aa298d28694bf5bd37d3e4912de8a2f5 (diff) | |
parent | 65b7d45852af3de557eaaf200471ff9b1a211970 (diff) | |
download | sandcrawler-e566ee1b4e134bfc06284cf77d8d1370df30d53f.tar.gz sandcrawler-e566ee1b4e134bfc06284cf77d8d1370df30d53f.zip |
Merge branch 'master' of git.archive.org:webgroup/sandcrawler
Diffstat (limited to 'pig/tests/files/papers_url_doi.cdx')
-rw-r--r-- | pig/tests/files/papers_url_doi.cdx | 7 |
1 files changed, 7 insertions, 0 deletions
diff --git a/pig/tests/files/papers_url_doi.cdx b/pig/tests/files/papers_url_doi.cdx new file mode 100644 index 0000000..1ad5792 --- /dev/null +++ b/pig/tests/files/papers_url_doi.cdx @@ -0,0 +1,7 @@ +#http://journals.ametsoc.org/doi/pdf/10.1175/2008BAMS2370.1 +#http://www.nejm.org:80/doi/pdf/10.1056/NEJMoa1013607 + +# should match 2: + +org,ametsoc,journals)/doi/pdf/10.1175/2008BAMS2370.1 20170706005950 http://mit.edu/file.pdf application/pdf 200 MQHD36X5MNZPWFNMD5LFOYZSFGCHUN3V - - 123 456 CRAWL/CRAWL.warc.gz +org,nejm,www)/doi/pdf/10.1056/NEJMoa1013607 20170706005950 http://mit.edu/file.pdf application/pdf 200 MQHD36X5MNZPWFNMD5LFOYZSFGCHUN3V - - 123 456 CRAWL/CRAWL.warc.gz |