Exporting Collection Data

View as MarkdownOpen in Claude

Once participants have completed your collection study, you can export all responses and uploaded files as a single ZIP archive.

Exports are generated asynchronously — the API responds immediately with a job ID, and the archive is built out-of-band. This keeps the request fast even for large collections with many uploaded files.

Workflow overview

1

Request an export by sending a POST request to the export endpoint. The API returns a job ID immediately.

3

Poll for completion by sending GET requests with your job ID until the status is complete or failed.

5

Download the ZIP archive from the presigned URL included in the complete response.

6

Work with the exported data — load responses.jsonl for analysis or extract files from the files/ directory.

Using the Prolific CLI

The Prolific CLI handles the full request, poll, and download flow in a single command:

prolific collection export <collection-id>

By default the archive is saved to <collection-id>-export-<YYYYMMDD-HHMMSS>.zip in the current directory. Use --output to specify a path:

prolific collection export <collection-id> --output ./my-export.zip

Requires the PROLIFIC_TOKEN environment variable and researcher access to the collection’s workspace.

The CLI currently only supports this all-in-one request/poll/download flow — it doesn’t yet support study_id/from/to filtering, or the listing and deleting operations described below. Use the API directly for those until CLI support lands.

Requesting an export

POST /api/v1/data-collection/collections/{collection_id}/export

No request body is required. By default, the export includes every response for the collection.

Filtering an export

Scope the export to a subset of responses with optional query parameters:

ParameterTypeDescription
study_idstringOnly include responses submitted under this Prolific Study ID.
fromISO 8601 datetimeOnly include responses with created_at on or after this datetime (inclusive).
toISO 8601 datetimeOnly include responses with created_at before this datetime (exclusive).
POST /api/v1/data-collection/collections/{collection_id}/export?study_id={study_id}&from=2026-01-01T00:00:00Z&to=2026-02-01T00:00:00Z

All three combine as AND — for example, study_id together with from/to exports only that study’s responses that were submitted within the given date range. from and to accept an offset (e.g. +01:00) as well as Z, and are normalized to UTC before comparison.

to must not be earlier than from — the API returns 400 Bad Request if it is.

If you use an offset instead of Z, URL-encode the + as %2B in the query string — a literal + is decoded as a space, which produces an invalid datetime and a 400. For example, 2026-01-01T00:00:00+01:00 should be sent as ...&from=2026-01-01T00:00:00%2B01:00.

Each distinct filter is tracked as its own export job: requesting a full export and a study_id-scoped export for the same collection produces two independent jobs (and two export_ids), and neither invalidates the other.

Responses

If a new export job is created, the API returns 202 Accepted:

{
"status": "generating",
"export_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
}

POST is idempotent only while a matching export is still generating — re-sending it for a collection with the exact same filter returns the existing job’s export_id rather than starting a duplicate. A completed or failed export is never reused, so re-sending POST after one finishes always starts a fresh export job, giving you a new snapshot of responses as of that request.

Each collection can have at most 10 export jobs at a time, across every filter combined. If you’re at that limit, the API returns 409 Conflict — delete an export you no longer need to free up a slot.

If the filter is invalid — to earlier than from, a malformed date, or a study_id with no responses in this collection — the API returns 400 Bad Request.

Polling for completion

Use the export_id from the POST response to check the status of your export job.

GET /api/v1/data-collection/collections/{collection_id}/export/{export_id}

Poll at a reasonable interval (every 3–5 seconds) until the status changes.

StatusMeaningNext step
generatingThe archive is still being builtContinue polling
completeThe archive is ready — url and expires_at are includedDownload the ZIP
failedGeneration failed or the archive was deletedRetry by sending POST again

Complete response

{
"status": "complete",
"url": "https://...",
"expires_at": "2026-03-20T10:30:00Z"
}

The url is a presigned HTTPS link valid for 1 hour. Re-poll the GET endpoint to receive a refreshed URL if it has expired.

Polling example

import time
import requests
def poll_export(collection_id, export_id, token, timeout=600):
headers = {"Authorization": f"Token {token}"}
deadline = time.time() + timeout
while time.time() < deadline:
r = requests.get(
f"https://api.prolific.com/api/v1/data-collection/collections/{collection_id}/export/{export_id}",
headers=headers,
)
r.raise_for_status()
data = r.json()
if data["status"] == "complete":
return data["url"]
if data["status"] == "failed":
raise RuntimeError(f"Export failed for collection {collection_id}")
time.sleep(3)
raise TimeoutError("Export did not complete within the timeout period")

Listing export jobs

Retrieve every export job requested for a collection, most recent first:

GET /api/v1/data-collection/collections/{collection_id}/export
[
{
"export_id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
"filter": null,
"status": "generating",
"created_at": "2026-03-19T11:00:00Z"
},
{
"export_id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
"filter": { "study_id": "645d8e4a1234567890abcdef" },
"status": "complete",
"created_at": "2026-03-19T10:30:00Z"
}
]

filter is null for an unfiltered export, otherwise the study_id/from/to the job was requested with. Manually deleted export jobs are excluded from this list. Use GET /collections/{collection_id}/export/{export_id} for a job’s full status, including its download URL once complete.

Deleting an export

Permanently delete an export job and, if it completed, its ZIP archive:

DELETE /api/v1/data-collection/collections/{collection_id}/export/{export_id}

Returns 204 No Content on success. Deletion is immediate and cannot be undone — to get the data again, request a new export.

An export that is still generating cannot be deleted — wait for it to reach complete or failed first, or the API returns 409 Conflict.

Deleting frees up a slot toward the 10-export-per-collection cap, and the job immediately disappears from GET /collections/{collection_id}/export.

Archive format

The downloaded ZIP contains the following structure:

collection-export-{collection_id}-{YYYYMMDDTHHMMSS}/
├── responses.jsonl
├── collection.json
├── README.md
└── files/
└── {submission_id}_{instruction_id}_{index}.{ext}

The files/ directory is only present if participants uploaded files. The README.md inside the archive contains a quick-start guide and a pandas example.

responses.jsonl

Each line is a JSON object representing one submission. Response values are keyed by instruction ID.

{
"submission_id": "sub-abc123",
"participant_id": "part-def456",
"collection_id": "0192a3b4-c5d6-7e8f-9a0b-1c2d3e4f5a6b",
"created_at": "2026-03-19T10:30:00Z",
"responses": {
"0192a3b4-e7f8-7a0b-1c2d-3e4f5a6b7c8d": {
"type": "free_text",
"description": "Briefly describe the skin condition shown in your image",
"value": "Small red patch on left forearm, slightly raised"
},
"0192a3b5-e3f4-7a5b-6c7d-9e0f1a2b3c4d": {
"type": "file_upload",
"description": "Upload a clear photo of the affected area",
"files": [
{
"name": "photo.jpg",
"path": "files/sub-abc123_0192a3b5-e3f4-7a5b-6c7d-9e0f1a2b3c4d_0.jpg"
}
]
}
}
}

The path for file uploads is relative to the archive root, so it can be used directly after extraction.

Response value shapes

Instruction typeValue fields
free_textvalue: string
free_text_with_unitvalue: string, unit: string
multiple_choicevalues: string[]
multiple_choice_with_free_textvalues: { option: string, explanation: string }[]
file_uploadfiles: { name: string, path: string }[]

collection.json

Collection metadata and a list of all instructions, useful for mapping instruction IDs to their descriptions and types.

{
"collection_id": "0192a3b4-c5d6-7e8f-9a0b-1c2d3e4f5a6b",
"name": "Skin Condition Image Collection",
"exported_at": "2026-03-19T10:30:00Z",
"instructions": [
{
"id": "0192a3b4-e7f8-7a0b-1c2d-3e4f5a6b7c8d",
"type": "free_text",
"description": "Briefly describe the skin condition shown in your image"
},
{
"id": "0192a3b5-e3f4-7a5b-6c7d-9e0f1a2b3c4d",
"type": "file_upload",
"description": "Upload a clear photo of the affected area"
}
]
}

Working with the exported data

Load responses with pandas

import pandas as pd
df = pd.read_json("responses.jsonl", lines=True)
print(df.head())

Extract uploaded files

import zipfile
with zipfile.ZipFile("export.zip") as z:
z.extractall("export/")
# Files are at: export/files/{submission_id}_{instruction_id}_{index}.{ext}

The submission_id prefix in each filename lets you match files back to their submission record in responses.jsonl.

Handle all response types

for record in df.itertuples():
for instruction_id, response in record.responses.items():
match response["type"]:
case "free_text" | "free_text_with_unit":
print(response["value"])
case "multiple_choice":
print(response["values"])
case "multiple_choice_with_free_text":
for v in response["values"]:
print(v["option"], v["explanation"])
case "file_upload":
for f in response["files"]:
print(f["path"])

Notes

  • Presigned URL expiry: download URLs are valid for 1 hour. Re-poll GET to receive a refreshed URL.
  • Retry on failure: a failed export can be retried by sending POST again.
  • Active responses only: deleted and no_submission responses are excluded from the export.
  • Export job cap: each collection can have at most 10 export jobs at a time, across every filter combined. Delete jobs you no longer need to free up slots.
  • Deletion is permanent: there’s no way to undo a DELETE — request a new export if you need the data again.

By using AI Task Builder, you agree to our AI Task Builder Terms.