Data standards
What a DataSteak file actually looks like when it lands - the format, the column conventions, how editions work and what the quality figures mean.
This page describes what you get, before you buy it. Every product on DataSteak follows the same conventions, so once you have loaded one you know how to load the rest.
The file
Every dataset is delivered as a CSV, UTF-8 encoded, comma separated, with a header row. One file per product. No zip, no folder structure, no second file to join.
The file is generated when you download it rather than sat on a shelf, so what you get is the current edition at the moment you ask for it.
There is a byte-order mark at the front. Excel will otherwise read a UTF-8 file as Windows-1252 and turn every accented character into mojibake, which for a UK company register means a few thousand names quietly corrupted. The BOM is there to stop that.
If you are reading the file in code rather than a spreadsheet, that BOM will otherwise show up glued to your first column name. In Python, open it with `encoding="utf-8-sig"` and it disappears. Most other languages have an equivalent.
One thing we change, and why
A CSV cell beginning `=`, `+`, `-`, `@`, a tab or a carriage return is treated as a formula by Excel, Sheets and LibreOffice. That is how a spreadsheet can be made to run something just by being opened, and a company register is full of names that start with punctuation.
So a cell starting with one of those characters, and which is not a number, gets a leading apostrophe. Negative numbers are untouched, because they are numbers.
Spreadsheets treat that apostrophe as a text marker and do not display it. A parser will see it, so if you are reading the file in code and a value looks like it has gained a leading quote, that is this and not a data error. Strip it if it matters to you.
Nothing else is altered. No trimming, no case changes, no silent substitutions.
Editions, not feeds
A dataset is published in editions. Between editions the contents do not change, and the edition date is on the product page, in the API response body, and in the `X-DataSteak-Edition` header.
Dates render as `1 Sep 2026`. Never as "today", never as "20 days ago". A relative label makes you do arithmetic to work out whether you already have this one.
Different datasets are republished on different schedules, because their sources are. Some monthly, some quarterly, some when the publisher gets to it. There is no site-wide refresh day, so do not build a job around a calendar you have guessed at. Read the edition date and let it tell you whether anything moved.
How done is it
Every data product carries one of four levels, and it means the same thing everywhere it appears.
- Raw — untouched source data, exactly as published.
- Rare — cleaned, typed and labelled. One source, nothing invented.
- Medium — joined across sources and quality checked.
- Well Done — finished intelligence, enriched and scored.
The line that matters is between Rare and Medium. Rare is a single source made usable. Medium is more than one source brought together. Neither adds anything that was not in the data, which is the difference between both of them and Well Done.
Columns
Names are lower case with underscores, and they match between the CSV and the API. `company_name`, `company_number`, `incorporated`, `sic_code`, `post_town`, `postcode`, `accounts_next_due`. The API returns the same columns with the same names and the same types as the file, so code written against one works against the other.
The product page is the exception, and only for display. The preview table there tidies names for reading, so `sic_code` appears as "Sic Code" and `accounts_next_due` as "Accounts Next Due". The file and the API always give you the raw form. Write your code against those, not against what the shop window shows.
Some columns are copied straight from the publisher and some are derived from the publisher's own values, a size band worked out from figures they published rather than a field they publish. Which of those you are getting is what the doneness level below is telling you.
SIC codes are five digits, resolved against the published Companies House SIC list rather than a subset of our own. A company can carry more than one; where a product includes them it says in the product description how many and in what order.
Two things worth knowing before you filter on them. A handful of records on the register carry a four-digit code that matches nothing, almost certainly leading zeros lost somewhere upstream. And a few thousand companies still sit on the 2003 SIC vocabulary, which does not map to the current one. Neither is something we invent a value for.
Dates are `YYYY-MM-DD` in the data, whatever the original source used. Several UK registers publish `DD/MM/YYYY` and at least one publishes `09/Jan/2018 - 00:00`; you will not see either of those in a file. Sort them as strings if you like, it works.
Numbers carry no thousands separators and no currency symbols. Formatting belongs to whatever you are building, not to the file.
A blank means the source had nothing there. It does not mean zero, and it does not mean we dropped it. Where a column is materially incomplete the product page publishes the figure rather than leaving you to find out after you have paid.
What the completeness figures mean
Each product page shows, per column, the percentage of rows that have a value. It is calculated across the whole cut, not a sample, and a value counts as present only if it is not blank after trimming whitespace.
The preview table on the product page is chosen to show full rows, deliberately, so the sample is not representative of completeness. The percentages directly beneath it are the honest number and they describe the whole file. If those two ever seem to disagree, trust the percentages.
Where a product says "active"
That is the register's own status, not a judgement of ours. We do not decide that a company has stopped trading; we pass on what the source says and name the field it came from.
The same applies to any status, rating or classification in a product. If it came from the publisher, it is theirs. If we derived it, the product page says so and the doneness level moves accordingly.
Why the products are cut the way they are
Datasets are published as filtered cuts by region, industry or characteristic rather than as one enormous file. Partly because a cut is more useful than a haystack. Partly for a reason that is less obvious:
A spreadsheet stops at 1,048,576 rows. Some of these sources are larger than that. A file you cannot open is not a product, so the catalogue is designed around cuts that fit in the tool most buyers will actually use. If you want something bigger than a cut, that is what the API and a bulk route are for, and you should ask.
Where a product is geographic, the boundary is stated on its page in terms you can check yourself. Postal areas rather than an administrative boundary we would then have to defend, so you can verify membership against a postcode rather than take our word for the shape.
Two licences, not one
The licence you buy from us governs what you may do with the product. Underneath it sits the original publisher's licence, which applies to you whatever we say. Several UK public registers are published under the Open Government Licence and require attribution when you publish anything derived from them. Others carry their own terms.
Every product names its source, so you can always tell which applies. The pricing and licensing page has ours, and the terms have the exact wording.
Looking for the data behind this?
Every dataset says where it came from, what is in it and how often it refreshes.