npx skills add ...
npx skills add catalyst-cooperative/agent-skills --skill pudl
Explore and understand PUDL energy data: discover which tables exist, look up column meanings and usage warnings, and load Parquet files from S3 or a local directory. No PUDL Python package required. Use this skill whenever a user asks what PUDL data contains, wants to understand a specific table or column, asks about data quality or limitations, needs help loading data into a notebook or script, or wants to know which table covers a topic like electricity generation, utility financials, fuel costs, power plant locations, emissions, capacity factors, FERC financial data, or EIA survey data. Also use when the user mentions PUDL, Catalyst Cooperative energy data, or any of the specific data sources PUDL ingests (EIA-860, EIA-861, EIA-923, FERC Form 1, FERC Form 714, FERC EQR, EPA CEMS, EPA CAMD, etc.).
npx skills add catalyst-cooperative/agent-skills --skill pudl
This skill is for data users who want to explore, understand, and load PUDL's public energy data products. It assumes no access to the PUDL Python package or source repository — only the publicly distributed data files and their metadata.
PUDL's primary outputs are Apache Parquet files, described by a Frictionless Data
Package descriptor. For generic descriptor-querying patterns (jq), use
the datapackage skill — this skill provides PUDL-specific knowledge layered on top.
Beyond the main Parquet outputs, PUDL also distributes raw per-form FERC Parquet data
(covering Forms 1/2/6/60/714, each with its own datapackage.json) and the FERC EQR
(partitioned Parquet, separate from the main build). These have different access
patterns and are not covered by the main Frictionless descriptor — see
Data Access for the full picture.
Every step below is inexpensive and should happen by default whenever it's relevant to the question at hand, not only when the user asks for it by name.
Locate the metadata — the primary PUDL descriptor (Parquet outputs) is at:
s3://pudl.catalyst.coop/nightly/pudl_parquet_datapackage.jsonhttps://s3.us-west-2.amazonaws.com/pudl.catalyst.coop/nightly/pudl_parquet_datapackage.jsonRaw per-form FERC data has its own datapackage.json in each form/era directory,
e.g. s3://pudl.catalyst.coop/nightly/ferc1_xbrl/datapackage.json and
s3://pudl.catalyst.coop/nightly/ferc1_dbf/datapackage.json — see
Raw per-form Parquet directories
for the full list.
The FERC EQR (Electric Quarterly Reports) is distributed separately due to its size, and only one version is publicly available at a time:
s3://pudl.catalyst.coop/ferceqr/ferceqr_parquet_datapackage.jsonhttps://s3.us-west-2.amazonaws.com/pudl.catalyst.coop/ferceqr/ferceqr_parquet_datapackage.jsonFor offline or development use, download all descriptors locally with:
This populates assets/cache/. The script is cache-aware — a cached file
younger than a day is reused with no network call, so it's safe to run this
every time you need a descriptor rather than checking assets/cache/ yourself
first. Pass --force to bypass the cache and refetch regardless of age (e.g. if
you suspect PUDL's schema changed today and need the very latest copy).
Raw input archives (for provenance) live at
s3://pudl.catalyst.coop/zenodo/<dataset>/<concrete-doi>/datapackage.json.
Prefer the cached S3 archive over the Zenodo website or API for raw metadata and
file access. The source docs page usually gives a concept DOI for the whole dataset
lineage; the S3 path uses a concrete DOI for one specific archived version. See
Data Quality and Context for details.
Query metadata selectively — use /datapackage skill patterns (jq)
to find relevant tables, read descriptions, and surface warnings.
For "does PUDL have data on X" questions, don't stop at a match you already
recognized by reputation — run a broader keyword sweep across relevant
description/code fields first (for FERC accounts, see
Cross-referencing FERC Form 1 and Form 2 schedules and accounts;
the same habit applies to other sources' core_*__codes_* tables). Flag it if
an answer came from recalled knowledge rather than the sweep.
Consult primary-source forms and instructions when metadata alone doesn't fully explain something — don't wait for the user to ask for these by name. See Data Sources: Blank forms and filer instructions.
Check table tier — see Data Quality and Context.
Prefer out_* tables; warn users about _core_* tables.
Check keys before joining tables — if the task combines a FERC-sourced table
with an EIA-sourced table (or any two tables at all), check schema.foreignKeys
on each first, and route utility/plant joins through utility_id_pudl /
plant_id_pudl, not through name-string matching. See
PUDL Datapackage Extensions: Joining PUDL tables.
Check methodology before implementation details — if the user is asking how
PUDL cleans, imputes, allocates, reconciles, estimates, or models data, read
Methodology first and fetch the relevant public
methodology page (append .md to the URL for your own reading — but when pointing
the user to it, give them the plain .html link) before looking at source code,
docstrings, or implementation details. Summarize the public methodology page and
point the user to it. Only dive into code-level implementation after the user has
seen that write-up or if no methodology page exists for the topic.
Load the data, efficiently — Loading data doesn't have to mean downloading an
entire table. SELECT ... LIMIT in DuckDB, pl.scan_parquet() with
.select()/.filter() before .collect() in polars, and a columns= argument in
pandas all push the selection down to the Parquet reader itself. Treat sampling and
down-selecting as the normal way to explore a table, not an optimization reserved
for when a file turns out to be huge. You should estimate a table's size before a
full, unfiltered load, and only load the full table if the job genuinely needs every
row; see Data Access for the loading patterns
themselves.
sources array (31 datasets, with short codes, names, licensing, and per-source
documentation links), and where to find and read each source's blank forms and
filer instructions; read when a user asks about a specific source dataset
(EIA-860, FERC Form 714, EPA CEMS, etc.) or needs documentation links, when
resolving a raw-archive S3 path and you need the short code and have to
distinguish between a concept-DOI and a concrete-DOI, or whenever interpreting what
a column, code, or schedule actually means.utility_id_pudl/plant_id_pudl; read
before querying description or other non-standard fields on a PUDL descriptor,
and before joining any two PUDL tables (for generic descriptor-querying mechanics,
use the datapackage skill instead)out_* vs core_* vs raw), warning types, and what each tier
means for analysis reliability; read when a user asks about data quality, when choosing
between table tiers, or when surfacing warnings before providing loading codeferc_electricity_accounts.json over reading this
fileferc1_schedules.json over reading this fileferc2_schedules.json over
reading this fileferc_accounts arrayLicense: All PUDL data is published under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. Users may freely use, share, and adapt the data with attribution to Catalyst Cooperative.
Citation: When a user asks how to cite PUDL, provide this reference:
Selvans, Z., Gosnell, C., Sharpe, A., Schira, Z., Lamb, K., Belfer, E., Xia, D., & Mazaitis, K. The Public Utility Data Liberation (PUDL) Project [Data set]. Catalyst Cooperative. https://doi.org/10.5281/zenodo.3653158
BibTeX:
The S3 bucket s3://pudl.catalyst.coop is free and publicly accessible — no
AWS credentials needed, and any ambient credentials (even invalid ones) should be
explicitly bypassed rather than assumed absent.
DuckDB, pandas, and polars each need explicit setup to query this bucket
reliably — see
Data Access: DuckDB and S3
(s3_url_style plus clearing S3 credential settings; applies through /query too)
and the pandas/polars sections below it (storage_options for anonymous access)
for why each is needed.
The Parquet path for a core PUDL output table is
s3://pudl.catalyst.coop/nightly/<table_name>.parquet. Raw per-form FERC tables
use a different path — see
Raw per-form Parquet directories.
Always surface usage warnings from the descriptor before providing loading code.
Methodology-first rule: if a public methodology page exists for the topic the user is asking about, use it before inspecting implementation details. Code-level explanations are a follow-up step, not the default first response.
Prefer out_* tables for analyst work. If a user asks about a topic without
specifying a table, search metadata for out_ tables first.
Use uv to install Python packages — prefer uv add <package> over
pip install <package>. uv is faster and installs into a virtual environment
rather than globally. Fall back to pip only if uv is not available
(command -v uv returns nothing) — and if you do, install into a project-local
virtual environment (create one with python -m venv .venv if none exists), not
the system/global Python. pip install --user is not a safe fallback either
— it still writes into the user's global user-site packages, shared across every
other project on their machine, rather than scoping the change to this task. If
the working directory already has its own environment manager (pixi, poetry, an
existing venv or conda env), install through that instead of introducing a second
one.
PUDL's datapackage descriptors extend the standard schema in several PUDL-specific
ways: RST-formatted, docstring-style descriptions, per-resource provenance metadata,
and a package-level unit registry. Read
PUDL Datapackage Extensions before writing
jq queries against description or other non-standard fields — it covers only
what's unique to PUDL; for generic descriptor-querying mechanics, use the
datapackage skill.
Prefer joining PUDL tables on ID columns over name-string columns
(utility_name_ferc1, utility_name_eia, plant names, etc.) — same-named
entities across FERC and EIA are not guaranteed to be the same company. Route
joins through utility_id_pudl / plant_id_pudl via the core_pudl__assn_*
crosswalk tables, checking schema.foreignKeys first. Name matching is a
legitimate fallback when no ID crosswalk is available, but treat its results as
unverified until spot-checked. See
PUDL Datapackage Extensions: Joining PUDL tables.
Both ferc1_schedules.json and ferc2_schedules.json share the same schema. Each
record has a ferc_accounts array with the account numbers that schedule references,
pre-extracted for direct lookup. Use description for topical keyword search; use
ferc_accounts for account-number cross-referencing.
Quick lookup patterns (jq):
Joining across both files (jq): load the accounts file with --slurpfile and use
INDEX() to build an account-number lookup, then join it against each schedule's
ferc_accounts array:
| User intent | Hand off to |
|---|---|
| Query datapackage.json metadata | /datapackage |
| Run SQL or NL queries against data | /query |