Batch Upload API
When to use it
Use the Batch Upload API when an integration needs to ingest many files from one job. A batch must include a shared bucket destination, or every item must provide its own destination.
Good fits:
- content migrations
- nightly sync jobs
- benchmark corpora
- generated documentation bundles
- bucket seeding scripts
Endpoints
Create the batch upload session in the destination bucket:
POST /v1/buckets/{bucket}/batches
Finalize after uploading accepted items:
POST /v1/buckets/{bucket}/batches/{batch_id}/finalize
Each batch accepts 1 to 100 files. {bucket} accepts a bucket id or slug.
POST /v1/knowledge/files:batch/upload-session (+ /{batch_id}/finalize) with bucket fields in the manifest.Create request
The manifest defines the batch and durable item ids. The files array supplies compact per-file upload metadata so Calypso can create one storage upload target per accepted item.
{
"manifest": {
"version": 1,
"batch_idempotency_key": "local-batch-001",
"items": [
{
"client_file_id": "contract_pdf",
"filename": "contract.pdf",
"title": "Customer Contract"
}
]
},
"files": [
{
"client_file_id": "contract_pdf",
"filename": "contract.pdf",
"content_type": "application/pdf",
"size_bytes": 184233
}
]
}
Every client_file_id must:
- be unique within the batch
- appear once in
manifest.items - appear once in
files - use only letters, numbers,
_,., or- - not start with reserved Firestore-style prefixes such as
__
Bucket routing
On the bucket-addressed path, every file in the batch lands in the path bucket — no bucket fields needed in the manifest. When one file needs different routing, put bucket_ids on that item (item-level fields still apply). On the legacy alias route, the old rules hold: a shared manifest-level destination or per-item bucket fields, else 400 bucket_required.
Create response
Calypso returns a batch id and one upload session per accepted item:
{
"batch_id": "batch_123",
"upload_strategy": "gcs_resumable",
"accepted": [
{
"client_file_id": "contract_pdf",
"session_id": "sess_123",
"upload_url": "https://storage.googleapis.com/...",
"expires_at": "2026-06-08T21:00:00Z"
}
],
"rejected": [],
"request_id": "req_current"
}
Upload every accepted file directly to its upload_url. Treat each URL as a short-lived bearer capability.
Finalize
curl -X POST "https://api.calypso.so/v1/buckets/rag1/batches/$BATCH_ID/finalize" \
-H "Authorization: Bearer sk-..." \
-H "Content-Type: application/json" \
-d '{"mode":"finalize_uploaded"}'
Finalize validates uploaded blobs item by item. Missing uploads stay pending. Already-finalized items are replayed, so retrying finalize is safe.
Each finalize call accumulates onto the results of earlier calls rather than replacing them, so finalizing a batch in several passes converges on one complete set of counters — including items rejected at creation time. A terminal item that fails during finalize is recorded as a failed batch item, so the failure is visible in batch polling and not only in the finalize response.
Items can also come back pending with the reason finalize_budget_exhausted: each finalize call runs under a server-side time budget, and items not reached in time are deferred rather than risking a timeout. This is normal for large batches — call finalize again and the remaining items complete (already-finalized items replay as no-ops).
Typical finalize response:
{
"batch_id": "batch_123",
"status": "finalized",
"finalized": [
{ "client_file_id": "contract_pdf", "session_id": "sess_123", "knowledge_id": "file_123", "task_id": "task_123" }
],
"pending": [],
"failed": [],
"replayed": [],
"request_id": "req_current"
}
Batch status
Poll:
GET /v1/knowledge/batches/{batch_id}?include_items=true
Batch statuses:
| Status | Meaning |
|---|---|
accepted | Files were stored and queued, but not all accepted items are terminal. |
indexing | At least one item is actively indexing. |
active | All accepted items are indexed. |
partially_active | Some files are active while others are queued, indexing, or failed. |
partially_failed | At least one item failed and not all accepted files failed. |
failed | All accepted files failed, or no files were accepted. |
While any item is still outstanding, polling deliberately withholds a terminal status (active, partially_failed, failed) — a batch never reports itself finished before every item has resolved, so it is safe to treat a terminal status as final.
Item statuses give the most useful debugging signal. Inspect item-level errors, knowledgeId, taskId, bucketSyncStatus, and bucketSync.
Query readiness
202 Accepted from finalize means the batch was durably accepted. It does not mean the files are query-ready.
Wait for each item's source to report ready: true (GET /v1/sources/{id}) — the one bit that covers both indexing and bucket sync. The item-level status, bucketSyncStatus, and bucketSync fields remain available for debugging.
Reliability rules
- Keep batch size at or below 100 files.
- Respect the 5 create requests per second per team rate limit.
- Use stable
batch_idempotency_keyvalues for retries. - Generate deterministic
client_file_idvalues when retrying the same local files. - Include a shared bucket destination, or include bucket fields on every item.
Next: