Skip to main content

Managed Datasets

An enterprise feature

Managed datasets are part of the enterprise edition. A community build registers no provider for them, serves none of their endpoints, and renders no record editor, so nothing on this page is reachable on one.

Every other dataset in the catalog got there by being captured: a scheduled query ran, and its results accumulated in the warehouse. A managed dataset works the other way round. Nothing feeds it: an administrator declares its shape, and then people type records into it through the product.

It exists for the reference data that never arrives on a wire, such as which stops have restrooms, which cameras are known to be miscalibrated, or which incidents were manually reclassified. That kind of data usually lives in a spreadsheet nothing can join against. This puts it in the warehouse beside everything else, where a query can reach it.

In the catalog a managed dataset carries the origin contributed and is marked Managed in the list. Where a dataset comes from covers how its page differs from a capture's.

Who can do what​

Three levels, and they are checked on the server rather than in the interface:

Can
Any signed-in memberRead the declaration, the records, and any record's revision history
A member of a writer groupAdd records as that group, and edit or retract the records that group owns
An administratorDeclare a dataset, change or delete one, and set which groups may write

A writer acts as a group rather than as themselves. Editing and retracting are limited to records whose own group is one you belong to and one the dataset still allows, so an admin who removes a group from the dataset closes editing on that group's existing records too, not only new ones.

Declaring one​

Admin → Managed Datasets lists every declaration on the instance with its provisioning state, and is where new ones are declared.

The admin console: three declarations, one ready, one provisioning and one failed with its error and a Retry You can also declare one from the catalog page of an existing dataset, which carries that dataset's shape over as a starting point; every prefilled value stays editable.

A declaration is a name, an optional description, the groups allowed to write, and a list of columns.

The id is derived, not typed​

The dataset's id becomes an unquoted ClickHouse table name, so it has to be a valid identifier. Rather than asking an admin to type one and then refusing it, the form derives the id from the name: lowercased, with anything outside a-z0-9_ replaced by an underscore, and a prefix added if the result would start with a digit. A hyphen is the natural thing to write in a name and exactly what the identifier rule forbids, so deriving the id avoids that refusal entirely.

The rule matches how captured columns are already named from query results, so a declared name and a captured one follow one rule rather than two that drift apart.

If the id collides with a declaration that already exists, the server answers 409 and the form reports it against the name, rather than adding a suffix and handing you an id you never chose.

Columns​

Each column has a name, a label shown in the table and the form, and one of seven types:

string, integer, float, boolean, date, timestamp, enum

A column cannot be named after one the record log or the underlying view already defines, and the declaration and its columns are written in a single transaction, so a rejected declaration leaves nothing half created.

Provisioning​

Declaring a dataset creates real warehouse objects, so a declaration has a state rather than existing immediately. The console shows that state (ready, provisioning, failed), prints the warehouse's own error on a failure, and offers a Retry on that row alone.

Writer groups are shown and chosen by name rather than by the group ids the declaration actually stores, so the column reads admin, default rather than a pair of numbers.

Delete is offered on every row, and the server refuses it for a declaration that has records. Empty ones delete normally.

Records​

On a contributed dataset's catalog page, a Records section sits below the header.

A managed dataset's page: the Records table, with edit and retract offered only on the rows this viewer's group owns

Note the last row in the screenshot. It was contributed by a group this viewer does not belong to, so it offers history and nothing else, while the three rows above it carry edit and retract as well.

The section carries its own heading. Without one, an empty dataset showed a lone Add record button under a schema table, with nothing saying what it would add a record to, and the editor was reported as unfindable.

The table​

One column per field the fetched records carry, labelled and ordered by the declaration rather than by whatever order the values happen to arrive in.

Paging is a cursor, and Next page replaces the page on screen rather than growing an ever longer list. The current page stays visible while the next one loads instead of flashing an empty table.

A captured dataset renders no Records section at all. It has no declaration to read, so the requests are never sent rather than sent and 404ed.

Adding and editing​

Add record opens a form generated from the declared columns, with each input typed to its column. It appears only when you belong to at least one group that may write to this dataset. Edit and retract icons appear per row, on the rows whose group you can write as.

Validation is per field: a value the declared ClickHouse type cannot hold is refused with the message placed against the input that caused it, rather than swept into one banner at the top of the form.

When permissions cannot be read

If the declaration read fails, Add, edit and retract are unavailable, and the panel says so rather than leaving the controls sitting there inert.

Retracting, and the history​

Records are an append-only log, so nothing here overwrites or deletes.

Retract appends a new revision at status retracted and drops the record out of the table. Its history is kept, and a writer can add a new record with the same values later. Retracting is not an undo, and there is no un-retract.

Every record carries its revision history, readable by any signed-in member, showing what each revision changed. An edit that would land on top of a revision the server has already moved past is refused with the current head returned, so the client can rebase rather than silently clobber someone else's edit.

Reaching the data from a query​

A managed dataset is a warehouse table like any other. Once provisioned it appears in the catalog, carries a schema, and can be queried, joined and put on a dashboard exactly as a captured dataset can, which is what it gains by living here rather than in a spreadsheet.