columns_cache_max_bytes_to_write_to_cache
Soft per-query threshold on the bytes a single query writes to the columns cache. The bytes written during the query are counted, and once the counter reaches this value, further cache writes for the rest of the query are skipped. This is an advisory threshold, not a hard cap: a reader accumulates the entries of the granules it has read and writes them to the cache in one batch, and the batch that crosses the threshold is stored in full before the counter is charged. So the actual amount written may exceed this value by up to the entries one reader accumulates between two writes - the columns it reads, one entry per stripe of about 65536 rows each - and, with several readers running at once, by that much per reader. The purpose is to keep a single large scan from displacing useful data from the cache, not to bound cache usage exactly. A value of0 means use half of the size limit the columns cache currently has: the configured columns_cache_size, or less while the cache is shrunk under memory pressure (see the ColumnsCacheSizeLimit metric).
columns_cache_max_estimated_bytes_to_write_to_cache
If the estimated size of the data a query reads fromMergeTree parts exceeds this value, writes to the columns cache are inhibited for the entire query. The estimate is made in uncompressed bytes, which is what the cache is charged for, from the size of the columns the query reads (including PREWHERE, mutation and patch-part columns) scaled to the selected mark ranges, and the query is charged for all of it before it reads anything. This keeps a single large scan from displacing useful data from the cache, and from copying data into the cache that cannot stay there.
A value of 0 means use half of the size limit the columns cache currently has. That is the configured columns_cache_size while the server has memory to spare, but the cache shrinks under memory pressure (see the ColumnsCacheSizeLimit metric), and the default budget shrinks with it. With the default columns_cache_size_ratio, half of the limit is the size of the probationary segment of the cache, so the data of a query that passes the gate can be cached completely in one pass.
The gate does not apply to a read that drops mark ranges while it runs, which is the case when use_indexes_refiner_in_read_pools is enabled: how many of the selected marks such a read really touches is decided only when each task is cut, so the estimate above would be an upper bound that charges marks the query never reads. For those reads the amount written is bounded by columns_cache_max_bytes_to_write_to_cache instead.