Integrations
Delivered where you already work
Every plan includes delivery to the destination you already use, so the data lands in your pipeline instead of another inbox. Pick a target below, or ask about one we haven't listed.
Delivery targets
Amazon S3
Data lands as files in a bucket you own, partitioned by product and date. Delivery uses a cross-account IAM role you grant us, so no long-lived keys change hands. New files follow a consistent naming pattern so your pipeline can pick them up without watching for arbitrary filenames.
Google Cloud Storage
Files are written to a bucket you control, using a service account scoped to that bucket only. Objects follow the same product/store/date partitioning as our other file-based targets. You can revoke the service account at any time without affecting delivery to your other destinations.
BigQuery
Rows are appended directly to a dataset in your project, authenticated through a service account with write access to that dataset. Each load includes a schema_version column so downstream queries can handle changes safely. Loads run on your chosen schedule, not continuously, to keep query costs predictable.
Snowflake
Data is loaded into a table in your account through a dedicated role and warehouse you provision for us. Loads are batched per run rather than streamed row by row, matching your refresh schedule. Credentials are scoped to that one role, with no access to the rest of your account.
PostgreSQL
We write to a schema in a database you provide, using a role limited to that schema. Each run upserts rows by SKU and timestamp rather than truncating the table, so a partial failure doesn't lose prior data. Connection details, including network access rules, are set up once during onboarding.
Google Sheets
Data refreshes into a sheet you share with a service account, which keeps formulas and formatting elsewhere in the spreadsheet untouched. This target suits smaller teams or a quick look at the data rather than a production pipeline. Row counts above a few thousand are better served by a file or database target.
CSV / JSON download
Files are made available for direct download, either pushed to a link we send or picked up from an agreed location. This is the simplest option for a one-off dataset or a trial sample, with no credentials to set up on your side. Ongoing feeds can start here and move to a managed destination later.
REST API webhook
We POST each run's results to an endpoint you provide, authenticated with a key or token you supply. This suits systems that need to react immediately, like a repricing engine, rather than polling for new files. Failed deliveries are retried on a backoff schedule, and we notify you if an endpoint stays unreachable.
Conventions
File naming
File-based deliveries follow product/store/YYYY-MM-DD/part-N.jsonl, so a pipeline can glob by date or store without parsing arbitrary filenames. Each part file is a complete, independently readable chunk, useful for splitting large deliveries across multiple files. The date reflects the run's scheduled date in UTC, not the time the file was written.
Schema versioning
Every delivery includes a schema_version column. Changes to a schema are additive; we add new fields rather than repurpose existing ones. Any field removal is announced at least 30 days ahead, so you have time to update anything that depends on it.
Refresh schedule
Weekly, daily and hourly refresh options run on a fixed time in UTC, not a rolling window, so you can predict exactly when new data arrives. If a scheduled run fails or runs late, we notify you rather than silently skipping it. Historical runs are not automatically backfilled unless agreed as part of the plan.
Need a destination we haven't listed?
Tell us what you use and we'll set it up as part of onboarding.