Skip to content

Data export

Requires a license with the data-export feature. See pricing.

Data export turns your content into a file. Pick the content types, the statuses and the format, and the export is built in the background. When it finishes you download the file or find it in your S3 bucket, and a schedule runs the same export on a timetable.

StepWhat happens
StartYou send the format and a filter. The answer is a job, at once.
BuildThe job moves through the stages queued, reading, writing, packaging (building a ZIP), storing and done.
CollectA finished job downloads from the Admin API, or lands in the S3 bucket you named.
RepeatA schedule starts the same export on a cron timetable, in UTC.

You need a content type post with a few entries, and a token in TOKEN for an admin or super_admin. The quickstart shows how to get one.

  1. Start a CSV export of post:

    Terminal window
    curl -X POST http://localhost:3001/api/admin/data-export/start \
    -H "Authorization: Bearer $TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"name": "posts-csv", "format": "csv", "filter": {"schemas": ["post"]}}'

    The answer is 202 with the job, "status": "pending" and "stage": "queued". Copy its id into JOB.

  2. Check it:

    Terminal window
    curl http://localhost:3001/api/admin/data-export/jobs/$JOB \
    -H "Authorization: Bearer $TOKEN"

    Within a moment it shows "status": "completed", "stage": "done" and rows_written.

  3. Download it:

    Terminal window
    curl http://localhost:3001/api/admin/data-export/jobs/$JOB/download \
    -H "Authorization: Bearer $TOKEN" -OJ

    You get posts-csv.csv.

  4. Schedule a nightly JSON export:

    Terminal window
    curl -X POST http://localhost:3001/api/admin/data-export/schedules \
    -H "Authorization: Bearer $TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"name": "Nightly", "cron_expression": "0 2 * * *", "enabled": true, "export_config": {"format": "json", "include_schemas": true}}'

    The answer is 201 with the schedule and its next_run_at, 02:00 UTC tonight. Copy its id into SCHEDULE.

  5. Pause it, and delete it when you are done:

    Terminal window
    curl -X PUT http://localhost:3001/api/admin/data-export/schedules/$SCHEDULE \
    -H "Authorization: Bearer $TOKEN" \
    -H "Content-Type: application/json" \
    -d '{"enabled": false}'
    curl -X DELETE http://localhost:3001/api/admin/data-export/schedules/$SCHEDULE \
    -H "Authorization: Bearer $TOKEN"

    The pause answers 200 without a next_run_at, and the delete answers 204.

  6. In the admin console, open Insight, then Data export. The export is under Exports, with a Download button. New export and New schedule do steps 1 and 4 from forms.

FormatWhat you get
jsonOne document: {"lyeve_export": 1, "exported_at": "...", "schemas": [...], "entries": [...]}. schemas appears only with include_schemas.
ndjsonOne entry per line.
yamlThe json document as YAML.
csvOne file per content type, <schema>.csv, with the entry columns (slug, title, status, published_at, created_at, updated_at, id) and then one column per field. A field named like an entry column, such as title, takes that column. Several content types download as a ZIP.
markdownEach entry as a heading and its body.
hugoMarkdown with TOML front matter.
templateYour own Go text/template, rendered once per entry. Send it in template.

In json, ndjson and yaml, an entry has id, schema, slug, title, status, published_at, created_at, updated_at, its fields under body, and meta when it has any. Tenant and author ids are left out.

In csv, numbers and booleans are written as they are and nested values as JSON. Text a spreadsheet would run as a formula gets a leading '.

Bulk import reads json, ndjson, yaml and csv exports back as they are, one content type at a time.

FieldMeaningDefault
formatRequired. One of the formats above.
nameA label for the job.export-<format>
filter.schemasSchemas to include. Every name must be a schema of the tenant.all
filter.statusesdraft, published, archived.all
filter.from, filter.toEntries updated at or after from and before to. RFC 3339.none
filter.slugOne entry by slug.none
filter.searchEntries whose title contains this text.none
filter.fieldsBody fields to keep.all
filter.order_byupdated_at, created_at or title.updated_at
filter.order_ascOldest or A to Z first.false
filter.limitMost entries. 0 means the default cap.10000
include_schemasAdd each schema's definition. Inside the document for json and yaml, as schemas.json beside the data otherwise, which makes the download a ZIP.false
include_mediaBundle the media files entries reference. The download becomes a ZIP. One file may be up to 5 GiB.false
flatten_jsonWrite a nested CSV value as key=value; key=value in its cell instead of JSON.false
s3{"bucket", "key", "region"}. Upload the result instead of keeping it for download. key is a prefix, and the extension is added.none

GET /api/admin/data-export/jobs/{id} answers the job with its progress: status, stage, rows_processed, rows_total, rows_written and file_size. status is pending, running, completed, failed or canceled. A job that fails or is canceled keeps the stage it stopped at, and error says why it failed.

Once status is completed, GET .../jobs/{id}/download answers the file as an attachment named after the job. An export sent to S3 answers 307 to a signed address in the bucket instead, so pass -L to curl. POST .../jobs/{id}/cancel stops a job that is still running.

A schedule holds an export's options under export_config, an optional s3 target and a cron_expression. It runs in UTC.

Terminal window
curl -X POST http://localhost:3001/api/admin/data-export/schedules \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{
"name": "Nightly backup",
"cron_expression": "0 2 * * *",
"enabled": true,
"export_config": {"format": "json", "include_schemas": true},
"s3": {"bucket": "lyeve-exports", "key": "nightly/content", "region": "eu-west-1"}
}'
  • cron_expression has five fields (minute, hour, day of month, month, day of week) or is @hourly, @daily, @weekly, @monthly, @yearly or @every <duration>. At most 128 characters.
  • Due schedules are checked every 30 seconds. With several replicas, each run starts exactly once.
  • A schedule that fell behind runs once and moves to its next time. Missed runs are not replayed.
  • next_run_at shows when it runs next and last_run_at when it last ran. Disabling a schedule clears next_run_at.
  • PUT changes only the fields you send.

In the admin console, New schedule offers Every hour, Every day at 02:00 UTC, Every Monday at 02:00 UTC, On the first of each month at 02:00 UTC or A cron expression of my own, and each schedule can be paused, resumed or deleted.

Every route takes an admin or super_admin and a signed-in session. An admin token cannot call them.

Export and schedule routes
MethodPathPurpose
POST/api/admin/data-export/startStart an export. Answers 202 with the job.
GET/api/admin/data-export/jobsList jobs. Takes status, limit (default 50, max 500) and offset. Answers {"jobs", "total", "count"}.
GET/api/admin/data-export/jobs/{id}One job with its progress.
POST/api/admin/data-export/jobs/{id}/cancelCancel a running job. Answers {"status": "canceled"}.
GET/api/admin/data-export/jobs/{id}/downloadDownload a completed export.
GET/api/admin/data-export/schedulesList schedules. Answers {"schedules", "total", "count"}.
POST/api/admin/data-export/schedulesCreate a schedule. Answers 201.
GET/api/admin/data-export/schedules/{id}One schedule.
PUT/api/admin/data-export/schedules/{id}Update a schedule.
DELETE/api/admin/data-export/schedules/{id}Delete a schedule. Answers 204.
VariableWhat it doesDefault
DATA_EXPORT_S3_ALLOWED_BUCKETSComma-separated buckets an export may upload to. my-org-* matches by prefix and * matches every bucket.empty, so every S3 export is refused
DATA_EXPORT_S3_ALLOWED_REGIONSComma-separated regions an upload may use.empty, so any region
DATA_EXPORT_MAX_CONCURRENTExports that may run at once.3

A schedule's S3 target is checked against the allowlist again every time it runs. If you set LYEVE_PLUGINS to choose which features start, include data-export in it. See licensing and tiers.

Flows can use data_export.start, which starts an export and returns the job, and data_export.status, which reads a job and, once it completes, its download route. A test run starts no export.

Without a license that carries data-export, the routes answer 404. A license that lapses while the instance runs makes them answer 402 payment_required.

Every error these routes return
StatusMessageCause
400invalid JSON bodyThe body is not JSON.
400format must be json, ndjson, yaml, csv, markdown, hugo, or templateMissing or unknown format.
400template is required for format=templateTemplate format without a template.
400filter.schemas names a schema this tenant does not define: <name>A schema name is wrong.
400S3 exports are not configured: ...DATA_EXPORT_S3_ALLOWED_BUCKETS is empty.
400S3 bucket "<name>" is not in the allowed listThe bucket is not allowed.
400S3 region "<name>" is not in the allowed listThe region is not allowed.
400name is required, cron_expression is requiredSchedule created without them.
400cron_expression must be five fields or a descriptor such as @dailyThe expression does not parse.
409job already in terminal state: <status>Cancel on a finished job.
409export not yet completed (status: <status>)Download before the job completes.
404export object not foundThe file is gone from storage.
503too many concurrent exports running, try again laterDATA_EXPORT_MAX_CONCURRENT exports are already running.
400S3 bucket is not a valid bucket nameThe bucket name breaks S3's rules.
402payment_requiredThe license lapsed while the instance ran.
404database errorNo job or schedule has that id.