Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

download_dv.py - Download datasets from Dataverse repositories

Description

This script downloads all files of a Dataverse dataset (Harvard Dataverse or any other Dataverse installation), preserving the dataset’s folder structure. It downloads each file separately. Dataverse caps the size of whole-dataset ZIP downloads and silently leaves out files beyond the cap, so a single ZIP download can be incomplete for large deposits.

Usage

python3 tools/download_dv.py IDENTIFIER [--server_url URL] [--output PATH] [--dry-run]
python3 tools/download_dv.py --jira-ticket KEY [--print-id]
python3 tools/download_dv.py IDENTIFIER --dir-name

Arguments

Examples

# Harvard Dataverse
python3 tools/download_dv.py doi:10.7910/DVN/T81OHQ

# Another Dataverse installation, by landing-page URL
python3 tools/download_dv.py "https://dataverse.nl/dataset.xhtml?persistentId=doi:10.34894/ABCDEF"

# Show what would be downloaded
python3 tools/download_dv.py https://doi.org/10.7910/DVN/T81OHQ --dry-run

# Identifier from Jira; prints only "dv-DVN-T81OHQ" on stdout
python3 tools/download_dv.py --jira-ticket AEAREP-9261 --print-id

What gets downloaded

Top-level ZIP files in the deposit (for example, Code.zip) are extracted later by 00_unpack_zip.sh in the pipeline.

Output Structure

Input DOI: doi:10.7910/DVN/ABC123
Output directory: ./dv-DVN-ABC123/

The directory name is dv- followed by the last two components of the DOI (dv-10.5064-F6ABCD for a two-part DOI such as 10.5064/F6ABCD).

Pipeline Use

The 1-populate-from-icpsr and w-big-populate-from-icpsr pipelines call this script when no openICPSR ID is set and no World Bank deposit was found. The Dataverse identifier comes from, in order:

  1. the DataverseID pipeline variable,

  2. the dataverse: field in config.yml,

  3. the Jira ticket’s “Replication package URL” field (for example, AEAREP-9261 has https://doi.org/10.7910/DVN/T81OHQ).

A Jira URL counts as a Dataverse deposit if it mentions dataverse, dataset.xhtml, or /DVN/, or if it is a DOI that resolves to a Dataverse dataset page. If the Jira URL is empty or is not a Dataverse deposit, the script exits with code 2 and the pipeline tries the Zenodo downloader next. When the identifier came from Jira, the pipeline writes it back to dataverse: in config.yml. Later steps then get the deposit directory name from --dir-name.

Large deposits exceed what the regular Download step can cache. Use w-big-populate-from-icpsr for them.

Exit codes

Git integration

When run in CI without --print-id, the script commits the downloaded files itself. With --print-id, the pipeline handles commits.

Requirements