Description¶
This script downloads all files of a Dataverse dataset (Harvard Dataverse or any other Dataverse installation), preserving the dataset’s folder structure. It downloads each file separately. Dataverse caps the size of whole-dataset ZIP downloads and silently leaves out files beyond the cap, so a single ZIP download can be incomplete for large deposits.
Usage¶
python3 tools/download_dv.py IDENTIFIER [--server_url URL] [--output PATH] [--dry-run]
python3 tools/download_dv.py --jira-ticket KEY [--print-id]
python3 tools/download_dv.py IDENTIFIER --dir-nameArguments¶
IDENTIFIER - DOI in any common form (
doi:10.7910/DVN/ABC123,10.7910/DVN/ABC123,https://doi.org/10.7910/DVN/ABC123) or a Dataverse landing-page URL (https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:...)--doi - Same as IDENTIFIER (kept for compatibility)
--server_url - Dataverse installation URL. If omitted, taken from a landing-page URL, or found by resolving the DOI; falls back to
https://dataverse.harvard.edu--output - Parent directory for the
dv-...output directory (default: current directory)--jira-ticket KEY - When no identifier is given, use the Jira ticket’s “Replication package URL” if it is a Dataverse deposit
--print-id - Send all progress output to stderr and print only the output directory name on stdout (for pipeline capture)
--dir-name - Print the output directory name for the identifier and exit, without any network access
--dry-run - List the files that would be downloaded
Examples¶
# Harvard Dataverse
python3 tools/download_dv.py doi:10.7910/DVN/T81OHQ
# Another Dataverse installation, by landing-page URL
python3 tools/download_dv.py "https://dataverse.nl/dataset.xhtml?persistentId=doi:10.34894/ABCDEF"
# Show what would be downloaded
python3 tools/download_dv.py https://doi.org/10.7910/DVN/T81OHQ --dry-run
# Identifier from Jira; prints only "dv-DVN-T81OHQ" on stdout
python3 tools/download_dv.py --jira-ticket AEAREP-9261 --print-idWhat gets downloaded¶
Every non-restricted file of the latest version, saved under its Dataverse folder (for example,
Data/raw/file.csv).Tabular files that Dataverse converted to
.tabare downloaded in their original format (.csv,.dta, ...), with their original names.Each file’s checksum and size are checked after the download. A mismatch or failed download counts as an error.
Restricted files cannot be downloaded without authentication. They are listed and skipped.
Top-level ZIP files in the deposit (for example, Code.zip) are extracted later by 00_unpack_zip.sh in the pipeline.
Output Structure¶
Input DOI: doi:10.7910/DVN/ABC123
Output directory: ./dv-DVN-ABC123/The directory name is dv- followed by the last two components of the DOI (dv-10.5064-F6ABCD for a two-part DOI such as 10.5064/F6ABCD).
Pipeline Use¶
The 1-populate-from-icpsr and w-big-populate-from-icpsr pipelines call this script when no openICPSR ID is set and no World Bank deposit was found. The Dataverse identifier comes from, in order:
the
DataverseIDpipeline variable,the
dataverse:field inconfig.yml,the Jira ticket’s “Replication package URL” field (for example, AEAREP-9261 has
https://doi.org/10.7910/DVN/T81OHQ).
A Jira URL counts as a Dataverse deposit if it mentions dataverse, dataset.xhtml, or /DVN/, or if it is a DOI that resolves to a Dataverse dataset page. If the Jira URL is empty or is not a Dataverse deposit, the script exits with code 2 and the pipeline tries the Zenodo downloader next. When the identifier came from Jira, the pipeline writes it back to dataverse: in config.yml. Later steps then get the deposit directory name from --dir-name.
Large deposits exceed what the regular Download step can cache. Use w-big-populate-from-icpsr for them.
Exit codes¶
0- Success1- Error (bad identifier, API error, failed or corrupt downloads)2- Not a Dataverse deposit (only with--jira-ticketand no identifier)
Git integration¶
When run in CI without --print-id, the script commits the downloaded files itself. With --print-id, the pipeline handles commits.
Requirements¶
Python 3 standard library only
--jira-ticketneedstools/jira_get_info.pyand Jira credentials (JIRA_USERNAME,JIRA_API_KEY)