xgboost (XGBOOST) dictionary trains an XGBoost gradient-boosted model, at load time, from a source table of training rows, then predicts a numeric target for any feature vector you pass in. The feature columns are the dictionary key and the single attribute is the target the model learns.
It is suited to tabular regression and binary classification where the features are numeric — for example forecasting a value from several measurements, or scoring rows against a learned target. Multiclass objectives are not supported (see Layout parameters).
The XGBoost integration is experimental. Enable it with the Only
enable_xgboost setting before creating an XGBOOST dictionary or calling predictXGBoost:CREATE DICTIONARY is supported: an XGBOOST dictionary defined in a server configuration file fails to load.predictXGBoost is the only way to query the dictionary: it takes the features as individual arguments, returns the prediction, and accepts additional prediction parameters. The dictionary holds a trained model rather than rows, so the generic dictionary interface — dictGet, dictHas and SELECT * FROM dict — is not supported and reports an error.
Quickstart
Here we train a regressor on the linear targety = 2*x1 + 3*x2.
1. Create a source table of training rows — the feature columns followed by the target:
XGBOOST layout — the feature columns are the key and y is the target attribute:
PRIMARY KEY (x1, x2) makes x1 and x2 the features. The target the model learns is y, inferred as the single column that is not part of the key; the parameters in LAYOUT are XGBoost hyperparameters (see Layout parameters).
4. Predict — predictXGBoost takes the features positionally and returns the prediction:
2*1 + 3*2 = 8, so the model’s prediction is close to 8.
How it works
Training (at load time). Each source row is a(features..., target) observation. When the dictionary loads, the source is read block by block and the model is then trained once over the whole set. Feature and target values are read as floats, so the key columns must be numeric and the target attribute floating-point (see Dictionary structure).
Predicting (at query time). To predict, the model takes the feature vector — in the same order as the key columns were declared — and runs it through the trained booster, returning a Float64. A NULL feature is passed to the model as a missing value, which XGBoost handles the way it learned to during training, so the prediction is never NULL. When every feature is a constant, the model is evaluated once per block instead of once per row.
The model is not persisted. It lives only in memory, for as long as the dictionary is loaded, and is trained again from the source on every load — including after a server restart.
Retraining the model. Because every load trains from scratch, SYSTEM RELOAD DICTIONARY retrains the model against the current contents of the source table:
LIFETIME also retrains, since a lifetime-triggered reload is an ordinary load. Use it to refresh the model periodically as the training data grows.
Dictionary structure
AnXGBOOST dictionary has a fixed shape:
- The
PRIMARY KEYis one or more columns of a native numeric type (integers and floats) — the features. At query time this “key” is the feature vector you pass in to predict, not a stored lookup key. The feature order is the key-column declaration order, andpredictXGBoostbinds its positional arguments to that order. - Alongside them, declare exactly one attribute of type
Float32orFloat64: the target the model learns. It is always inferred as the single column that is not part of the feature key — there is no parameter to name it, and it is an error to declare more than one attribute.
Layout parameters
Only the parameters listed below are accepted; any other name fails the load, so typos are caught when the model trains rather than being silently ignored.num_iterations is handled by ClickHouse (see its description); every other parameter is forwarded to the XGBoost booster unchanged, as a string, and takes XGBoost’s own default and value range — see the XGBoost parameter reference.
Parameter names are case-insensitive, and each may be given only once. Values must be a positive integer, a float, or a quoted string: a negative literal is rejected by the dictionary DDL itself, before the layout sees it, so a negative
seed cannot be expressed.
For example, a dictionary that also sets the step size eta:
Prediction parameters
predictXGBoost accepts an optional trailing constant Map of XGBoost prediction parameters, after the features, built with map:
XGBoosterPredictFromDMatrix. Only the keys below are accepted; any other key fails the query. Every parameter is an integer or a boolean, so the Map values must be an integer type.
Notes
-
Computational dictionary semantics. This is a computational dictionary: it holds a trained model, not rows, and
predictXGBoostis the only way to query it. The generic dictionary interface is not supported and reports an error:dictGet(there is no stored attribute to look up — the “key” is a feature vector to predict from),dictHas(no keys are stored),SELECT * FROM dictand joining the dictionary as a table. BecausepredictXGBoostis the only entry point, theenable_xgboostsetting must be enabled for every prediction. -
Numeric columns only. Every feature (key) column must be a native numeric type and the target attribute must be
Float32orFloat64. Values are read as floats during training and prediction. The feature arguments ofpredictXGBoostmay also beNullable, see How it works. -
system.dictionariesreports no stored items. The dictionary trains a model instead of storing rows, soelement_countis0, as it is for adirectdictionary, andbytes_allocatedis0too: the trained model belongs to XGBoost, which does not report how much memory it holds.query_countandfound_ratecount the rows passed to the model bypredictXGBoost; a call whose features are all constant is evaluated once per block, so it counts one row per block rather than one per row. -
A failed reload keeps the previous model. If retraining fails — the source table is gone, its schema changed, a hyperparameter is no longer accepted — the dictionary does not start failing predictions. It keeps serving the last model that trained successfully, and records the error instead. Compare
last_successful_update_timewithlast_exceptioninsystem.dictionariesto tell whether the model still reflects the current source data: -
XGBoost’s messages go to the server log. They are written under the
XGBoostlogger, at the level XGBoost gives them, and are attached to the query that trains or predicts, so they also appear insystem.text_log. Raiseverbosityto see more of them. -
Feature order matters.
predictXGBoostbinds its positional feature arguments to the key columns in declaration order, and the number of feature arguments must match the number of key columns.