Exporting Collection Data
Once participants have completed your collection study, you can export all responses and uploaded files as a single ZIP archive.
Exports are generated asynchronously — the API responds immediately with a job ID, and the archive is built out-of-band. This keeps the request fast even for large collections with many uploaded files.
Workflow overview
Request an export by sending a POST request to the export endpoint. The API returns a job ID immediately.
Using the Prolific CLI
The Prolific CLI handles the full request, poll, and download flow in a single command:
By default the archive is saved to <collection-id>-export-<YYYYMMDD-HHMMSS>.zip in the current directory. Use --output to specify a path:
Requires the PROLIFIC_TOKEN environment variable and researcher access to the collection’s workspace.
Requesting an export
No request body is required. By default, the export includes every response for the collection.
Filtering an export
Scope the export to a subset of responses with optional query parameters:
All three combine as AND — for example, study_id together with from/to exports only that study’s responses that were submitted within the given date range. from and to accept an offset (e.g. +01:00) as well as Z, and are normalized to UTC before comparison.
to must not be earlier than from — the API returns 400 Bad Request if it
is.
If you use an offset instead of Z, URL-encode the + as %2B in the query
string — a literal + is decoded as a space, which produces an invalid
datetime and a 400. For example, 2026-01-01T00:00:00+01:00 should be sent
as ...&from=2026-01-01T00:00:00%2B01:00.
Each distinct filter is tracked as its own export job: requesting a full export and a study_id-scoped export for the same collection produces two independent jobs (and two export_ids), and neither invalidates the other.
Responses
If a new export job is created, the API returns 202 Accepted:
POST is idempotent only while a matching export is still generating —
re-sending it for a collection with the exact same filter returns the existing
job’s export_id rather than starting a duplicate. A completed or failed
export is never reused, so re-sending POST after one finishes always starts
a fresh export job, giving you a new snapshot of responses as of that request.
Each collection can have at most 10 export jobs at a time, across every filter combined. If you’re at that limit, the API returns 409 Conflict — delete an export you no longer need to free up a slot.
If the filter is invalid — to earlier than from, a malformed date, or a study_id with no responses in this collection — the API returns 400 Bad Request.
Polling for completion
Use the export_id from the POST response to check the status of your export job.
Poll at a reasonable interval (every 3–5 seconds) until the status changes.
Complete response
The url is a presigned HTTPS link valid for 1 hour. Re-poll the GET endpoint to receive a refreshed URL if it has expired.
Polling example
Listing export jobs
Retrieve every export job requested for a collection, most recent first:
filter is null for an unfiltered export, otherwise the study_id/from/to the job was requested with. Manually deleted export jobs are excluded from this list. Use GET /collections/{collection_id}/export/{export_id} for a job’s full status, including its download URL once complete.
Deleting an export
Permanently delete an export job and, if it completed, its ZIP archive:
Returns 204 No Content on success. Deletion is immediate and cannot be undone — to get the data again, request a new export.
An export that is still generating cannot be deleted — wait for it to reach
complete or failed first, or the API returns 409 Conflict.
Deleting frees up a slot toward the 10-export-per-collection cap, and the job immediately disappears from GET /collections/{collection_id}/export.
Archive format
The downloaded ZIP contains the following structure:
The files/ directory is only present if participants uploaded files. The README.md inside the archive contains a quick-start guide and a pandas example.
responses.jsonl
Each line is a JSON object representing one submission. Response values are keyed by instruction ID.
The path for file uploads is relative to the archive root, so it can be used directly after extraction.
Response value shapes
collection.json
Collection metadata and a list of all instructions, useful for mapping instruction IDs to their descriptions and types.
Working with the exported data
Load responses with pandas
Extract uploaded files
The submission_id prefix in each filename lets you match files back to their submission record in responses.jsonl.
Handle all response types
Notes
- Presigned URL expiry: download URLs are valid for 1 hour. Re-poll
GETto receive a refreshed URL. - Retry on failure: a
failedexport can be retried by sendingPOSTagain. - Active responses only: deleted and
no_submissionresponses are excluded from the export. - Export job cap: each collection can have at most 10 export jobs at a time, across every filter combined. Delete jobs you no longer need to free up slots.
- Deletion is permanent: there’s no way to undo a
DELETE— request a new export if you need the data again.
By using AI Task Builder, you agree to our AI Task Builder Terms.