Configuration Drift Lives Between Files, So Lint Env Files, YAML, TOML and CI Together
.env.example says DATABASE_URL=postgres://localhost:5432/app. The compose file the team uses for local work mounts a volume and sets DATABASE_URL=sqlite:///data/app.db. The CI workflow exports DB_URL, a name the code stopped reading two refactors ago. Production has a secret called DATABASE_URL that nobody has looked at since the last rotation. Each of those files is valid, and every linter you could point at them passes them. The app still behaves differently on three machines, and the first person to find out is whoever is on call.
Drift is the gap between files that were each right when someone wrote them. It doesn’t live inside any one file, so a tool that checks files one at a time can’t find it. The version worth building parses everything (env files, YAML, JSON, TOML, compose files, Kubernetes manifests, CI workflows, Terraform variables), scans the source for the keys the code reads, and builds one graph: this key is defined here, read there, in these environments. Contradictions then become queries against the graph. This post is about what that tool would look like; the command and file names below are placeholders.
Each Linter Stays Inside Its Own File Type
dotenv-linter checks the style and consistency of env files. yamllint checks YAML syntax and style. kube-linter and kubeconform check Kubernetes manifests against best practices and schemas. conftest runs OPA policies over structured config, so a determined team can write cross-file rules in Rego; what it lacks is any notion of a variable the code reads. CUE can vet existing YAML and JSON against a schema, and Pkl, Apple’s configuration language (open-sourced in February 2024), renders typed modules into the formats your tools expect. Those two give real cross-file guarantees, at the price of writing the schemas by hand and, for the full benefit, moving your source of truth into a new language.
Frameworks offer the cheapest check there is. A settings class (pydantic-settings in Python, envalid in Node) fails at startup when a required variable is missing. That’s a good habit, and it only fires in the environment the app starts in, at deploy time, after the bad config has shipped.
There’s a second way to kill drift: generate every file from one source so they can’t disagree. Preconfiguration does that for coding-agent setup files, writing the file each platform wants from one preconfig.yaml and checking existing ones with preconfig check. Generation wins when you can adopt the new source of truth. A lint pass works on the files as they are today, which makes it the easier first step in an old repo, and the two approaches complement each other.
A Graph of Keys
Extraction produces two kinds of facts. Definitions say a key gets a value in some environment: a line in .env, an environment: entry in compose, an env: block in a workflow, a Kubernetes ConfigMap, a variable in a Terraform file. References say something consumes a key: ${VAR} interpolation in compose (with its ${VAR:-default} fallback form), ${{ secrets.X }} in a workflow, a secretKeyRef in a pod spec. Then the source scan finds reads: process.env.X and process.env["X"], os.getenv("X"), os.environ["X"], os.Getenv("X"), plus whatever settings accessor a repo configures. Use a real parser for each language instead of a regex, so commented-out code and string literals don’t count.
An environment is a named set of sources, declared in a small file, because no tool can guess that .env plus the compose file is “dev” and the manifests under deploy/ are “prod”:
[environments.dev]
files = [".env", "docker-compose.yml"]
[environments.ci]
files = [".github/workflows/*.yml"]
[environments.prod]
files = ["deploy/k8s/*.yaml"]
secret_names = "secrets/prod.names" # names only, exported by a trusted job
With every key placed in the graph, the rules are small queries. A key the code reads with no default and an environment never defines. A key that’s defined and never read. A key whose URL scheme or value type differs between environments, such as postgres:// here and sqlite:/// there. A value like DEBUG=false read through a plain truthiness check, since os.getenv("DEBUG") returns the string 'false', which is truthy. A name that looks secret (ending in _KEY, _TOKEN or _PASSWORD) with a value that doesn’t look like a placeholder. A key the code reads that’s missing from .env.example.
Secret scanners such as gitleaks and TruffleHog look for credentials by pattern and entropy, and they do it better than a drift tool would, so the secret rule stays a cheap name-and-placeholder check. Credentials sitting in config files are a standing item among mobile security mistakes too, which is a reason to catch them at review rather than in a bundle someone downloads. Output goes to the terminal as text and to SARIF, the format GitHub code scanning reads, so a finding appears on the pull request at the line that caused it:
$ cfgdrift scan
ERROR read-but-undefined STRIPE_WEBHOOK_SECRET
read at app/billing/webhooks.py:31 (no default)
defined in: dev (.env:14), ci (test.yml:22)
missing from: prod
ERROR scheme-conflict DATABASE_URL
dev .env:3 postgres://localhost:5432/app
dev docker-compose.yml:18 sqlite:///data/app.db
app/db.py:12 picks its driver from this scheme
WARN never-read DB_URL
defined in ci (test.yml:19); no read found in the 3 languages scanned
3 findings (2 errors, 1 warning); SARIF written to cfgdrift.sarif
The Parts That Resist
Dynamic key names come first. os.environ[f"{service}_URL"], or a loop over a list of names, can’t be resolved statically. Treat the read as a pattern (*_URL), match it against definitions, and report “unresolved dynamic read” instead of a false “never read” for every key the pattern might cover. Let people declare the pattern in a comment when the tool can’t see it.
Overlays are worse. Kustomize overlays, Helm values merged in layers and per-branch preview environments mean the effective value of a key sits in no single file; it exists only after rendering. Running helm template or kustomize build first and linting the output is the accurate answer, and it ties the linter to those tools and to values only known at deploy time. A first version should point an environment at rendered output and say so. Preview environments are the worst case, because the platform decides which settings are isolated per branch and which stay shared (Cloudflare Worker previews show one version of that split), and that decision lives in the platform rather than in a file.
Secrets you can’t read are a design problem. The question “does production define STRIPE_WEBHOOK_SECRET” has an answer in the secret store, but the store won’t show values and shouldn’t. GitHub’s API lists a repository’s secret names without values, and cloud secret managers separate listing from reading with permissions. Checking prod still needs credentials that can list secrets, and handing those to a CI job is a security decision. The design that holds up is a separate step: a trusted job exports names to a file, the linter reads the file, and the linter never holds a credential.
Monorepos change what a key is. The same name in six services is six keys, so identity has to be the pair of service and name, where a service is a directory with its own manifest, Dockerfile or deploy file. Without that, “defined but never read” flags every PORT in the repo. It flags a few real keys, too, that something outside the repo reads: the platform reads PORT, many libraries read NODE_ENV. A built-in allowlist for those saves the first week of annoyance.
False positives decide whether the tool stays installed. Rules need severity per repo and an ignore entry that carries a reason. A baseline file records existing findings so a legacy repo doesn’t open with hundreds of errors on day one; baseline what exists, fail only on what’s new. A linter that skips this gets uninstalled by Friday.
What the First Version Does and Refuses
Version 0.1 reads env files, compose files and GitHub Actions workflows. It scans source in three languages (JavaScript and TypeScript, Python, Go), ships five rules (read but undefined, defined but never read, scheme or type conflict, plaintext secret by name, missing from .env.example) and writes text plus SARIF. It refuses to manage secrets, since it never reads a value from a secret store, and it refuses to rewrite config. Autofix sounds friendly until it edits production’s file. A finding with a file and a line is enough.
“Defined but never read” is the config version of finding code, tables and flags nobody uses, and the tool shares a spine with a dependency feature map: parse everything, link definitions to uses, report where the graph looks wrong.
Nobody reads four config files side by side during review. A linter will.