← Learning Hub
Developer

Using the DataSteak API

13 Sep 2026 · 5 min read

How to pull a dataset into your own systems over HTTP, including authentication, paging, and the one header that saves you most of the work.

Every dataset on DataSteak can be delivered as a CSV download. It can also be pulled directly by your own systems over HTTP, which is what this page is about.

API access is arranged per dataset against your account rather than bought from the site. Ask us for the dataset you want and we will turn it on. You then issue your own key from your dashboard.

What it is, and what it is not

It returns the same data as the file. Same columns, same names, same types. What changes is who does the fetching.

It is not a live feed. Every dataset is published in editions, and the API serves whatever the current edition is. Between editions the data does not change.

Different datasets are republished on different schedules, because their sources are. Some are monthly, some are quarterly, some arrive when the publisher feels like it. So do not build your job around a calendar you have guessed at: read the edition date and let it tell you. That is the whole reason it is on every response.

Your key

Create one in your dashboard under Business account, API Keys. Only an Owner or an Admin can.

The key is shown once, at the moment you create it. We store only a hash of it, so we cannot show it to you again and neither can anyone who gets hold of our database. If you lose it, revoke it and make another. Revoking takes effect immediately.

Send it as a header. Keys are never accepted in the URL, because a credential in a query string ends up in server logs, browser history and referer headers, which is three copies nobody meant to make.

curl "https://datasteak.co.uk/api/dataset.php?slug=birmingham-companies&limit=100" \
  -H "Authorization: Bearer ds_your_key_here"

If something between you and us strips the `Authorization` header, which some proxies do, send `X-API-Key: ds_...` instead.

What can this key reach

Call the endpoint with no `slug` and it tells you, so you never have to keep a list in sync by hand:

{
  "datasets": [
    {
      "slug": "birmingham-companies",
      "title": "Birmingham Active Companies",
      "expires_at": "2026-10-13 18:26:29",
      "url": "https://datasteak.co.uk/api/dataset.php?slug=birmingham-companies"
    }
  ],
  "count": 1
}

A valid key that reaches nothing is a normal state, not a fault. It means the key works and no dataset has been turned on for you yet.

Reading a dataset

Pass `slug`, and optionally `limit` and `offset`. `limit` defaults to 100 and caps at 1000.

{
  "dataset": { "slug": "birmingham-companies",
               "title": "Birmingham Active Companies",
               "rows": 149843,
               "source": "Companies House" },
  "edition": { "updated": "2026-09-01" },
  "access":  { "expires_at": "2026-10-13 18:26:29" },
  "page":    { "limit": 100, "offset": 0,
               "returned": 100, "next_offset": 100 },
  "data":    [ { "company_name": "...", "company_number": "...", ... } ]
}

Keep following `next_offset` until it comes back `null`. It is deliberately `null` rather than a number on the last page, so a loop terminates on something it cannot mistake for a valid offset.

The header that saves you the work

Every response carries the edition date, in the body and as a header:

X-DataSteak-Edition: 2026-09-01
X-DataSteak-Access-Expires: 2026-10-13 18:26:29

Store the edition date after each pull and compare it before the next one. If it has not moved, there is nothing new and you can skip the fetch entirely.

This matters more than it looks, and more the less often a dataset changes. A job polling daily against something republished monthly finds news on one day in thirty; against something quarterly, one day in ninety. Checking the edition first turns all the rest into cheap requests, for you and for us. It also means the same code works across datasets on completely different schedules, without you tracking any of them.

`X-DataSteak-Access-Expires` is there so your integration can warn you before access lapses rather than discovering it as a sudden `403`.

Limits

Sixty requests per key per ten minutes. Exceed it and you get `429` with a `Retry-After` header. The limit is per key rather than per address, so a scheduler on a shifting IP is fine, and two customers behind one corporate gateway do not throttle each other.

Paging is intended for reading slices. Pulling an entire large dataset page by page re-runs the query for every page, which is slow for you and expensive for us. If you need a whole dataset regularly, tell us and we will sort out a better route than 150 requests.

What the responses mean

  • `401` the key is missing, unknown or revoked. One message covers all three on purpose: you know which applies to you, and anyone else is guessing.
  • `403` the key is fine, that dataset is not turned on for your account.
  • `404` no dataset with that slug.
  • `429` rate limited. Wait, then retry.
  • `503` the endpoint is up but could not read that dataset. Usually temporary; tell us if it persists.

What you may do with it

The same licence as the file you would have downloaded. The data is yours to use in your own products, analysis and client work. Redistributing it as-is is what the licence restricts. The pricing and licensing page has the detail, and the terms have the exact wording.

There is a second licence underneath ours: the one the original publisher attaches to the data. It differs by source, and it applies to you whatever we say. Several UK public registers are published under the Open Government Licence, which requires attribution when you publish anything derived from them; others carry their own terms.

Each response names the source it came from in the `source` field, and the product page states it too, so you can always tell which applies to the rows in front of you.

Looking for the data behind this?

Every dataset says where it came from, what is in it and how often it refreshes.