npx skills add ...
npx skills add microsoft/skills --skill kql
KQL language expertise for writing correct, efficient Kusto Query Language queries. Covers syntax gotchas, join patterns, dynamic types, datetime pitfalls, regex patterns, serialization, memory management, result-size discipline, and advanced functions (geo, vector, graph). USE THIS SKILL whenever writing, debugging, or reviewing KQL queries — even simple ones — because the gotchas section prevents the most common errors that waste tool calls and cause expensive retry cascades. Trigger on: KQL, Kusto, ADX, Azure Data Explorer, Fabric Real-Time Intelligence, EventHouse, Log Analytics, log analysis, data exploration, time series, anomaly detection, summarize, where clause, join, extend, project, let statement, parse operator, extract function, any mention of pipe-forward query syntax.
npx skills add microsoft/skills --skill kql
Try it yourself: All
✅examples in this skill can be run against the public help cluster:https://help.kusto.windows.net, databaseSamples(containsStormEvents,SimpleGraph_Nodes/Edges,nyc_taxi, and more).
Kusto Query Language (KQL) is a pipe-forward query language for exploring data. It is the native query language for Azure Data Explorer (ADX), Microsoft Fabric Real-Time Intelligence (EventHouse), Azure Monitor Log Analytics, Microsoft Sentinel, and other Microsoft data services.
KQL queries are a chain of operators separated by |. Data flows left to right:
KQL has two execution planes:
| Plane | Starts with | Examples |
|---|---|---|
| Query | Table name, let, print, datatable | StormEvents | where State == "TEXAS" |
| Management | .show, .create, .set, .drop, .alter | .show tables, .show table T schema |
Management commands can be followed by query operators (the output is tabular), but the entire request runs on the management plane. You cannot start with a query and pipe into a management command.
When in doubt: if the first token starts with ., it's a management command. For a full catalog of schema exploration commands, see references/discovery-queries.md.
KQL's dynamic type is flexible but strict in certain contexts. A common mistake is using a dynamic column in summarize by, order by, or join on without casting.
The rule: Any time you use a dynamic-typed column in by, on, or order by, wrap it in an explicit cast.
Self-correction: When you see "is of a 'dynamic' type" in an error, add tostring(), tolong(), or todouble().
KQL joins have constraints that differ from SQL.
KQL join conditions support only ==. No <, >, !=, or function calls in join predicates.
For range joins, pre-bin values: | extend bin_val = bin(Value, 100), then join on bin_val. Note: values near bin boundaries may land in adjacent bins — consider checking neighboring bins or overlapping the range for precision.
Both sides of a join on clause must reference column entities only — not expressions, not aggregates.
Always check cardinality before joining tables with >10K rows. A cross-join explosion was the source of the single E_RUNAWAY_QUERY error (25K × 195 = potential 4.8M rows).
KQL handles regex natively — no need for Python.
extract_all gotchaUnlike Python's re.findall(), KQL's extract_all requires capturing groups in the regex:
| Function | Use case | Example |
|---|---|---|
extract(regex, group, source) | Single match | extract(@"User '([^']+)'", 1, Msg) |
extract_all(regex, source) | All matches (needs ()) | extract_all(@"(\w+)", Text) |
parse | Structured extraction | parse Msg with * "User '" Sender "' sent" * |
matches regex | Boolean filter | where Url matches regex @"^https?://" |
replace_regex | Find and replace | replace_regex(Text, @"\s+", " ") |
Window functions need serialized (ordered) input.
Functions requiring serialization: row_number(), row_cumsum(), prev(), next(), row_window_session().
The most common memory error. Caused by scanning too much data without pre-filtering.
| count to understand table size| where before | summarize — filter time range, partition key, or category firstdcount() on high-cardinality columns without pre-filteringmaterialize() for subqueries referenced multiple timesE_LOW_MEMORY_CONDITIONThe query touched too much data. Your options:
| where filters (time range, partition key)by columns in summarize| sample 10000 for exploratory work instead of full scansE_RUNAWAY_QUERYA join or aggregation produced too many output rows. Check join cardinality — one or both sides is too large.
Large results slow down analysis. Prevention:
| Query type | Safeguard |
|---|---|
| Exploratory | Always end with | take 10 or | take 20 |
| Aggregation | Use | top 20 by ... not unbounded summarize |
| Wide rows (vectors, JSON) | | project only needed columns |
make_list() / make_set() | Avoid on high-cardinality groups (produces huge cells) |
| Unknown size | Run | count first |
The vector trap: Tables with embedding columns (1536-dim float arrays) produce ~30KB per row. Even | take 20 yields 600KB. Always | project away vector columns unless you specifically need them.
KQL sometimes requires explicit casts when comparing computed string values — even when both sides are already strings.
This is most common with computed values from geo_point_to_s2cell() and strcat() comparisons. When in doubt, cast with tostring().
KQL handles these natively — no need for Python:
For detailed examples and patterns, consult references/advanced-patterns.md.
When you encounter an error, look it up here before retrying:
| Error message contains | Likely cause | Fix |
|---|---|---|
is of a 'dynamic' type | Dynamic column in by/on/order by | Wrap in tostring()/tolong() |
Only equality is allowed | Range predicate in join condition | Pre-bucket with S2/H3 cells or bin() |
extractall(): matching groups | Missing () in regex | Add (): @"(\w+)" not @"\w+" |
row set must be serialized | Window function on unsorted data | Add | serialize or | order by before it |
Cannot compare values of types string and string | Computed string comparison | Add tostring() on both sides |
Failed to resolve column named 'X' | Wrong column name or wrong table | Run .show table T schema to check column names |
E_LOW_MEMORY_CONDITION | Query touched too much data | Add | where filters, reduce time range, break into steps |
E_RUNAWAY_QUERY | Join/aggregation produced too many rows | Check cardinality before joining; add pre-filters |
for each left attribute, right attribute | Join on clause incomplete | Use explicit form: on $left.X == $right.Y |
needs to be bracketed | Reserved word used as identifier | Use ['keyword'] syntax |
plugin doesn't exist | Unavailable plugin on this cluster | Fall back to equivalent function or Python |
Expected string literal in datetime() | Bare integer in datetime literal | Use datetime(2024-01-01) not datetime(2024) |
Unexpected token after by | Complex expression in summarize by-clause | extend the expression first, then summarize by the column |
not recognized / unknown operator | Operator not available on this engine | Check operator support; try equivalent (order by = sort by) |
Datetime literals are a common source of errors. A wrong literal format can cascade into completely different approaches instead of fixing the small issue.
| Function | Purpose | Example |
|---|---|---|
bin(ts, 1h) | Round down to bucket boundary | bin(Timestamp, 1d) |
startofmonth(ts) | First day of month | startofmonth(Timestamp) |
datetime_part("hour", ts) | Extract component | datetime_part("year", Timestamp) |
format_datetime(ts, fmt) | Format as string | format_datetime(Timestamp, "yyyy-MM") |
ago(1d) | Relative time | where Timestamp > ago(1d) |
between(a .. b) | Range filter (inclusive) | where Timestamp between (datetime(2024-01-01) .. datetime(2024-01-31T23:59:59)) |
todatetime(str) | Parse string → datetime | todatetime("2024-01-15T10:30:00Z") |
totimespan(str) | Parse string → timespan | totimespan("01:30:00") |
KQL has subtle differences from SQL syntax.
| Entity | Convention | Example |
|---|---|---|
| Tables | UpperCamelCase | StormEvents, NetworkLogs |
| Columns | UpperCamelCase | StartTime, EventType |
Variables (let) | snake_case | let filtered_events = ... |
| Built-in functions | snake_case | format_bytes(), geo_distance_2points() |
| Stored functions | UpperCamelCase | .create function GetTopUsers |
Both sort by and order by work identically in KQL — they are aliases. Use whichever you prefer, but be consistent.
When a first KQL query fails, the temptation is to abandon the entire approach and try something completely different. The correct response is almost always to fix the specific error, not change strategy.
Rules for error recovery:
parse operator is often simpler than extract() for structured text:Before running any KQL query, mentally check:
| where before any | summarize| take N or | top Nby/on/order by is wrappedextract_all patterns have () around what you want to capturedcount() before joining| project to drop unneeded columnsdatetime(2024-01-01) not datetime(2024) or bare integers| extend first, then | summarize by the computed column