Snowflake Queue Time: How to Balance Your Query Workloads

Snowflake Queue Time | Keebo Blog

Nobody budgets for Snowflake queuing. But it’s often a sign that your warehouses are running less efficiently than they need to. And it’s easy to miss unless you know where to look. 

Queuing in Snowflake is often the first visible sign that a single, poorly optimized warehouse has stopped being a localized annoyance and turned into an avoidable line item. Over time, that waste compounds, dragging down the entire pipeline.

I’ve written before about a hygiene framework for Snowflake pipeline efficiency, but queuing deserves its own conversation. Let’s walk through what it is, what it signals about performance, and how to address its root causes. 

What Is Queuing in Snowflake?

The mechanics of Snowflake queuing are straightforward. When you send more concurrent queries to a warehouse than it can handle, Snowflake queues those queries until resources free up. That wait time is recorded in the queued_overload_time column of QUERY_HISTORY, a place many teams never think to check.

Snowflake bills per second while a warehouse is running. Auto-suspend only kicks in when the warehouse goes idle. Since a warehouse stuck in a queue is never idle, every minute your queries spend waiting in line is a minute of fully billed compute.

In other words, Snowflake queuing is a leading indicator of compute time you’re buying, with nothing happening.

What Does Queuing Tell You About Query Costs?

An overloaded warehouse does more than make queries wait in a queue. It also forces the queries that are running to fight for the same CPU and memory. Under pressure, queries that usually fit in memory start spilling to local disk and then to remote storage, and every spill makes execution drastically slower. Slower execution, in turn, creates more contention for the next query. 

Think about one bloated, inefficient query. Maybe it scans ten times more partitions than it needs. Or, maybe an AI assistant wrote it with a re-evaluated CTE instead of a materialized one. Rather than waste its own compute, it inflates the cost of every other query unlucky enough to share the warehouse. Basically, the inefficiency of one workload gets added to the invoice of every other team.

What Does Queuing Signal for the Rest of Your Stack?

The inefficiency behind queuing spills over the Snowflake boundary, acting as a force multiplier for waste across your stack. This happens in several ways.

First, you can see it in your dependency chains. When an overloaded warehouse holds a model in queue for two minutes, every downstream model starts late, causing those delays and inefficiencies to accumulate down the entire DAG. Meanwhile, the warehouses across your account wake up to do nothing, still billing their 60-second minimums and idle toward auto-suspend. An overloaded warehouse A burns money on warehouses B and C.

Second, there is the external tax, namely from your cloud provider. Your orchestrator and serverless compute both bill by duration. AWS is explicit about this in its own cost guidance: when code makes a blocking call, you’re billed for the time it waits on the response. The Airflow project considers waiting workers expensive enough that it built deferrable operators specifically to free worker slots during external polls. If your code makes a blocking call to a busy warehouse, you pay AWS to wait, pay for warehouse time while you queue, and pay both again when the retry fires. Unless your entire stack is fully asynchronous (submit, disconnect, poll), you are paying for the same idle time at least twice.

And when this compounded inefficiency causes you to miss your SLA targets, it gets treated like a data incident. Monte Carlo’s data quality survey found organizations average 61 data incidents a month. How many first showed up as a simple queue, an early sign of an overloaded warehouse pushing a load window past its deadline? Engineers spend hours debugging a pipeline that was actually fine; the real problem was the warehouse it ran on.

What Are the Indirect Costs Behind Snowflake Queuing?

Two other places this cost lands don’t appear on a monthly statement, but they’re arguably more expensive: 

Should I Size Up My Warehouse to Fix Queuing?

Eventually, the drag on performance becomes impossible to ignore. The easy fix: more compute. You size up or convert to a multi-cluster warehouse that spins up extra clusters when queries queue. I’ve made this trade myself. It always feels smart in the moment.

But this quick fix doesn’t solve the core issue of an inefficient workload. The fix was just provisioning more compute to muscle through it, which can burn through as many as double the credits. Essentially, you pay more to run unoptimized queries faster. 

How Do I Reduce Queuing in Snowflake?

So if sizing up your warehouse is the wrong answer, what is? Here are three steps I’d recommend to reduce queuing by attacking the root cause of poorly written queries. 

1. Put a Number on Your Queue Time

Sum queued_overload_time by warehouse over the last 30 days and divide it by total execution time:

SELECT warehouse_name,   ROUND(SUM(queued_overload_time) / 1000, 0) AS queued_seconds,   ROUND(SUM(execution_time) / 1000, 0) AS execution_seconds,   ROUND(100 * SUM(queued_overload_time) / NULLIF(SUM(execution_time), 0), 1) AS queue_tax_pct FROM snowflake.account_usage.query_history WHERE start_time > dateadd('days', -30, current_date) GROUP BY warehouse_name ORDER BY queue_tax_pct DESC;

If you see more than a few percent on anything serving humans or SLA-bound pipelines, it’s a sign you’re leaving money on the table. Little’s Law guarantees the relationship you will find: the wait you measure is a direct function of arrival rate and the concurrency you provisioned.

2. Isolate Before You Upsize

Find out what is at the front of the line. Moving one rogue, heavy-duty job to its own right-sized warehouse beats a size bump. It costs less, not more. Performance isolation between warehouses is not a workaround, either. It is a designed property of the architecture, described in Snowflake’s own SIGMOD paper.

3. Make the Response Autonomous

Queue pressure fluctuates hour to hour, and workloads drift month to month. The warehouse size that was right in January is wrong by March, and no human re-tunes sizes, cluster counts, and suspend settings continuously. That’s why manual fixes always regress. 

This is the gap my team at Keebo is focused on: autonomous optimization that watches the queue signal around the clock, rightsizes in real time, and keeps learning as your workload changes, so the fix you make today is still the right fix six months from now.

The signal is already in your account, timestamped to the millisecond in SNOWFLAKE.ACCOUNT_USAGE.QUERY_HISTORY.

Stop paying for the inefficiency behind your queue time. See what my team at Keebo is doing to resolve it and reduce its impact on your bill.