
One source of truth for DNS, DHCP and IPAM
A reproducible self-hosted DDI proof of concept where NetBox drives Technitium DNS, Kea DHCP and the Proxmox VE SDN through event-driven Windmill flows. Full Docker Compose stack and configuration included.
Every lab I run catches the same disease. One IP address ends up living in four places: a DNS record, a DHCP reservation, a VLAN on the Proxmox host, and a spreadsheet that stopped being true months ago.
Nothing breaks until you delete a machine. Then three of those four keep pointing at it.
The fix I settled on: one source of truth, everything else generated from it. All self-hosted, nothing leaving the rack. What follows is enough to rebuild the whole thing from an empty directory.
On this page
Architecture
source of truth orchestrator enforced state
┌────────────────────┐ ┌──────────────────┐ ┌────────────────────┐
│ NetBox │ │ Windmill │ ──▶ │ Technitium DNS │
│ │ │ │ ├────────────────────┤
│ prefixes [dhcp] │ ──▶ │ flow: sync DNS │ ──▶ │ Kea DHCP4 │
│ IPs (dns_name) │ │ flow: sync DHCP │ ├────────────────────┤
│ IPs (mac_address) │ │ flow: sync SDN │ ──▶ │ Proxmox VE SDN │
│ VLANs │ │ │ │ (VNets per VLAN) │
└────────────────────┘ └──────────────────┘ └────────────────────┘
▲ ▲ ▲
│ │ │
humans webhook on change hourly reconcile| Role | What runs it | Why this one |
|---|---|---|
| IPAM | NetBox | Real data model, proper API, changelog on every object |
| Orchestration | Windmill | Flows as versioned Python steps, plus a UI to run them |
| DNS | Technitium | Full HTTP API, sane zone handling, lightweight |
| DHCP | Kea | Config is JSON and reloadable over a control socket |
| Network overlay | Proxmox VE SDN | VLANs already exist in NetBox, so they may as well match |
Three Docker networks, and the split matters:
netbox-net: NetBox, its Postgres and its Redis.ddns-net: DNS, DHCP and Proxmox.windmill-net: Windmill and its own database.
Windmill is the only container on all three. It is therefore the only thing that can reach both the source of truth and the targets.
Ports 53 and 67 are not published on the host, deliberately. The resolver
and the lease server answer inside ddns-net and nowhere else. That one
decision kills an entire class of “why is the Docker host resolving through the
lab” problems.
Why Windmill, after two attempts that were not?
Windmill is the third orchestrator this POC ran on. The first two were not mistakes. They were the shortest path to understanding what the job needed.
First attempt: a Python container on cron
A sync/ directory, three scripts, and an entrypoint that looped over them:
SYNC_INTERVAL = int(os.environ.get("SYNC_INTERVAL", "60"))
def job() -> None:
run_sync("sync_dns")
run_sync("sync_dhcp")
run_sync("sync_proxmox")
job()
schedule.every(SYNC_INTERVAL).seconds.do(job)
while True:
schedule.run_pending()
time.sleep(1)One SYNC_INTERVAL in .env, 60 seconds by default. It worked. The sync logic
inside it is essentially what still runs today.
What it could not do was tell me anything:
- Every run is a full poll. Nothing changed for six hours? It still hit every endpoint 360 times. And one interval had to serve three very different jobs.
- Failures scroll past.
run_synccatches, logs and carries on. That part is right: one broken target must not stop the other two. But the only trace is a line indocker logs, gone at the next restart. “Did the DHCP sync succeed at 14:03, and what did it push?” has no answer.
The rest follows: nothing to run on demand, and read, compute and apply are one
function call. Debugging meant adding a print and waiting another 60 seconds.
Second attempt: n8n
Next I moved the same logic into n8n, driven by NetBox webhooks. The workflows
went into the repository as n8n/workflows/*.json.
Event-driven was the right call, and it stayed. n8n as infrastructure as code was not.
Look at what a workflow file actually is. Here is one node, as committed:
{
"id": "22222222-2222-4222-8222-222222222222",
"name": "Lire Netbox",
"type": "n8n-nodes-base.code",
"typeVersion": 2,
"position": [550, 300],
"parameters": {
"mode": "runOnceForAllItems",
"jsCode": "const NETBOX_URL = 'http://netbox:8080';\nconst NETBOX_AUTH = '__NETBOX_AUTH__';\n\nasync function netboxGetAll(path) {\n const items = [];\n let url = `${NETBOX_URL}/api/${path}...
}
}Every real line of code sits inside one JSON string, newlines escaped. That single detail poisons everything:
- It cannot be reviewed. Change one character and
git diffshows one enormous modified line, and no tooling reaches what is inside it: no highlighting, no linter, no types. The editor sees a string. - Version control holds UI state.
"position": [550, 300]is where the box sits on the canvas. Move a box in the browser, get a diff.
Editing in the browser is pleasant. Right up to the moment two sources of truth exist, and you are diffing generated JSON by hand.
What Windmill changed
Windmill keeps what n8n got right: event-driven, a DAG (directed acyclic graph) you can look at, per-step results in a UI. It drops the rest, because a flow step is a normal Python file on disk:
windmill/scripts/sync_dhcp/
├── read_netbox.py # read
├── compute_config.py # compute
└── apply_kea.py # applysetup.py reads those files and pushes them into Windmill.
| Concern | Cron container | n8n | Windmill |
|---|---|---|---|
| Trigger | Fixed interval only | Webhook + schedule + manual | Webhook + schedule + manual |
| Logic in Git | Real .py files | Escaped JSON strings | Real .py files |
| Diff you can review | Yes | No | Yes |
| Per-step results | No | Yes | Yes |
| Run history | docker logs | Yes | Yes |
| Secrets | Env vars | Placeholder substitution | Typed secret variables |
| RBAC | None | Enterprise edition | Users, groups, folders |
| Maintenance burden | One image, no state | A database, plus workflows to re-export | A dedicated PostgreSQL |
What that buys:
- The code is linted and diffed like any other Python.
- Each step’s input, output and duration survives the run.
- A flow can be replayed from the UI while debugging.
- Secrets are variables, not strings interpolated into source at boot.
The honest caveat: this is not free. Windmill wants its own PostgreSQL, which is one more database to run and back up. If your sync is a single script with no branching and nobody else ever reads it, the cron container was fine. I would not talk you out of it.
Repository layout
.
├── compose.yaml
├── .env
├── config/
│ ├── kea/
│ │ ├── kea-dhcp4.conf # subnet4 is empty, the sync fills it
│ │ └── kea-ctrl-agent.conf # Unix socket exposed over HTTP
│ ├── netbox/
│ │ └── configuration.py
│ └── proxmox/
│ └── init_sdn.py # one-shot: API token + SDN zone
└── windmill/
├── setup.py # deploys variables, flows, schedules, webhooks
└── scripts/
├── sync_dns/{read_netbox,sync_technitium}.py
├── sync_dhcp/{read_netbox,compute_config,apply_kea}.py
└── sync_proxmox/{read_netbox_vlans,sync_proxmox_sdn}.pyThe Docker Compose stack
The whole POC lives in one compose.yaml. It is long, so here are the parts
that carry a decision. NetBox and its dependencies first:
name: ddi-stack
networks:
netbox-net:
ddns-net:
windmill-net:
services:
postgres:
image: postgres:16-alpine
restart: unless-stopped
networks: [netbox-net]
environment:
POSTGRES_DB: netbox
POSTGRES_USER: netbox
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
volumes:
- postgres-data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U netbox"]
interval: 10s
retries: 5
netbox:
image: netboxcommunity/netbox:latest
restart: unless-stopped
networks: [netbox-net, ddns-net]
ports:
- "8000:8080"
environment:
SECRET_KEY: ${NETBOX_SECRET_KEY}
NETBOX_TOKEN_PEPPER: ${NETBOX_TOKEN_PEPPER}
DB_HOST: postgres
REDIS_HOST: redis
REDIS_CACHE_HOST: redis
REDIS_CACHE_DATABASE: "1"
volumes:
- ./config/netbox/configuration.py:/etc/netbox/config/configuration.py:ro
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
healthcheck:
test: ["CMD-SHELL", "curl -sf http://localhost:8080/login/ || exit 1"]
interval: 30s
retries: 10
start_period: 120sTwo things worth copying.
start_period: 120s on the NetBox healthcheck. First boot runs migrations
and takes minutes. Without it, Compose calls the container unhealthy and
everything downstream gives up.
DNS and DHCP:
technitium:
image: technitium/dns-server:latest
networks: [ddns-net]
ports:
- "5380:5380" # web UI and API only, never 53
environment:
DNS_SERVER_ADMIN_PASSWORD: ${TECHNITIUM_PASSWORD}
volumes:
- technitium-data:/etc/dns
kea:
image: jonasal/kea-dhcp4:2
command: -c /etc/kea/kea-dhcp4.conf
networks: [ddns-net]
volumes:
- kea-leases:/kea/leases
- kea-sockets:/kea/sockets # shared with the control agent
- ./config/kea:/etc/kea
healthcheck:
test: ["CMD-SHELL", "test -S /kea/sockets/kea-dhcp4-ctrl.sock || exit 1"]
interval: 15s
start_period: 20s
kea-ctrl-agent:
image: jonasal/kea-ctrl-agent:2
command: -c /etc/kea/kea-ctrl-agent.conf
networks: [ddns-net]
ports:
- "8080:8000"
volumes:
- kea-sockets:/kea/sockets
- ./config/kea:/etc/kea
depends_on:
kea: { condition: service_healthy }Kea is split into two containers deliberately:
┌──────────────┐ Unix socket ┌────────────────┐ HTTP :8000
│ kea-dhcp4 │◀────────────────▶│ kea-ctrl-agent │◀────────────── the sync
│ speaks DHCP │ kea-sockets vol │ speaks HTTP │
└──────────────┘ └────────────────┘kea-dhcp4speaks DHCP and exposes a Unix control socket.kea-ctrl-agentis the only thing that turns that socket into HTTP.- They share it through the
kea-socketsvolume. - The healthcheck on the socket file stops the agent starting before there is anything to talk to.
Windmill and its setup job:
windmill:
image: ghcr.io/windmill-labs/windmill:main
networks: [windmill-net, netbox-net, ddns-net] # the only bridge
ports:
- "8300:8000"
environment:
DATABASE_URL: postgresql://windmill:${WINDMILL_DB_PASSWORD}@windmill-db/windmill
BASE_INTERNAL_URL: http://windmill:8000
SUPERADMIN_SECRET: ${WINDMILL_SUPERADMIN_SECRET}
NUM_WORKERS: "1"
depends_on:
windmill-db: { condition: service_healthy }
windmill-setup:
image: python:3.12-slim
restart: on-failure # retries until Windmill answers
command: sh -c "pip install -q requests && python /setup.py"
networks: [windmill-net, netbox-net, ddns-net]
volumes:
- ./windmill/scripts:/windmill-scripts:ro
- ./windmill/setup.py:/setup.py:ro
environment:
WINDMILL_URL: http://windmill:8000
NETBOX_URL: http://netbox:8080
NETBOX_TOKEN: ${NETBOX_TOKEN}
DNS_ALLOWED_ZONES: ${DNS_ALLOWED_ZONES:-}
KEA_DNS_SERVERS: ${KEA_DNS_SERVERS:-9.9.9.9, 149.112.112.112}
DHCP_TAG: ${DHCP_TAG:-dhcp}
depends_on:
windmill: { condition: service_healthy }
netbox: { condition: service_healthy }restart: on-failure on the setup container is the cheap way to handle
ordering. It exits non-zero while Windmill boots, Compose restarts it, it
eventually succeeds. No wait loop to write.
Kea, configured to be overwritten
The Kea config on disk is deliberately almost empty:
{
"Dhcp4": {
"interfaces-config": {
"interfaces": ["*"],
"dhcp-socket-type": "udp"
},
"control-socket": {
"socket-type": "unix",
"socket-name": "/kea/sockets/kea-dhcp4-ctrl.sock"
},
"lease-database": {
"type": "memfile",
"persist": true,
"name": "/kea/leases/dhcp4.leases"
},
"hooks-libraries": [
{ "library": "/usr/local/lib/kea/hooks/libdhcp_lease_cmds.so" }
],
"valid-lifetime": 4000,
"subnet4": []
}
}"subnet4": [] is the point. No subnet is ever written by hand. The sync
owns that array outright, and that is what makes “delete the prefix in NetBox”
actually remove the subnet.
libdhcp_lease_cmds.so is loaded so lease commands work over the control
channel. The agent just bridges the socket:
{
"Control-agent": {
"http-host": "0.0.0.0",
"http-port": 8000,
"control-sockets": {
"dhcp4": {
"socket-type": "unix",
"socket-name": "/kea/sockets/kea-dhcp4-ctrl.sock"
}
}
}
}Bringing it up
cp .env.example .env
# fill in NETBOX_SECRET_KEY, NETBOX_TOKEN_PEPPER, POSTGRES_PASSWORD,
# TECHNITIUM_PASSWORD, WINDMILL_SUPERADMIN_SECRET, WINDMILL_PASSWORD
docker compose up -d
docker compose logs -f netbox # migrations, 2 to 3 minutes on first bootCreate the NetBox superuser and an API token:
docker compose exec -e DJANGO_SUPERUSER_PASSWORD=change-me netbox \
/opt/netbox/venv/bin/python /opt/netbox/netbox/manage.py \
createsuperuser --username admin --email admin@lab.example --noinput
docker compose exec netbox /opt/netbox/venv/bin/python \
/opt/netbox/netbox/manage.py shell -c "
from users.models import Token
from django.contrib.auth import get_user_model
u = get_user_model().objects.get(username='admin')
print('NETBOX_TOKEN=' + Token.objects.create(user=u).key)
"Put that token in .env as NETBOX_TOKEN, then replay the setup job:
docker compose run --rm windmill-setupThat one command does everything:
- creates the Windmill workspace,
- pushes the variables, secret and plain,
- deploys the three flows from
windmill/scripts/, - registers the hourly schedules,
- creates the NetBox event rules.
Nothing is clicked in a UI. It is idempotent, so re-running it updates instead of duplicating.
Declaring the data in NetBox
Four fields carry everything.
Create the dhcp tag once (Customization → Tags), and the mac_address custom
field on IPAM > IP address (Customization → Custom Fields, type Text). Then:
| NetBox object | Value | Tag | Becomes |
|---|---|---|---|
| Prefix | 10.60.20.0/24 | dhcp | Kea subnet, router option .1 |
| IP range | .100 to .200 | dhcp | Dynamic pool inside that subnet |
| IP address | 10.60.20.10 | - | A record when dns_name is set |
| IP address | + mac_address | - | Fixed reservation in Kea |
Creating a host that gets both a record and a reservation:
curl -s -X POST http://localhost:8000/api/ipam/ip-addresses/ \
-H "Authorization: Token ${NETBOX_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"address": "10.60.20.10/24",
"status": "active",
"dns_name": "web01.lab.example",
"custom_fields": {"mac_address": "aa:bb:cc:00:11:22"}
}'Nobody fills in the default gateway. The flow computes it as the first host of the network. One less value to get wrong.
The flows, step by step
This is the heart of the project, and it is humbler than it sounds: make a few APIs talk to each other, on a trigger or on a schedule. No wheel is reinvented here. Data is translated from one API into another, and the translation is kept true.
Every flow is the same three-step DAG. That shape is the whole debugging story:
trigger (webhook │ schedule │ manual run)
│
▼
┌──────────────┐
│ read NetBox │ failed here? the source gave me nothing usable
└──────┬───────┘
▼
┌──────────────┐
│ compute │ failed here? I built the wrong config
└──────┬───────┘
▼
┌──────────────┐
│ apply │ failed here? the target rejected it
└──────────────┘One glance at which box is red tells you where to look. That is worth more than it sounds.
Sync DNS: read NetBox
Step one is a paginated read, nothing else:
NETBOX_URL = "http://netbox:8080"
def netbox_get_all(path: str, auth: str) -> list:
items, sep = [], "&" if "?" in path else "?"
url = f"{NETBOX_URL}/api/{path}{sep}limit=200"
while url:
r = requests.get(url, headers={"Authorization": auth})
r.raise_for_status()
data = r.json()
items.extend(data.get("results", []))
url = data.get("next")
return items
def main():
auth = f"Token {wmill.get_variable('u/admin/netbox_token')}"
desired = {}
for ip in netbox_get_all("ipam/ip-addresses/", auth):
if not ip.get("dns_name"):
continue
desired[ip["dns_name"].strip().rstrip(".")] = ip["address"].split("/")[0]
return {"desired": desired, "dns_records": len(desired)}Sync DNS: apply to Technitium
Step two picks the zone by longest suffix match. So
web01.dev.lab.example lands in dev.lab.example when that zone exists, and in
lab.example otherwise:
def find_zone(fqdn: str, zones: set) -> str | None:
best = None
for z in zones:
if fqdn == z or fqdn.endswith(f".{z}"):
if best is None or len(z) > len(best):
best = z
return bestThe write is an upsert, so replaying the sync is a no-op:
requests.get(f"{TECHNITIUM_URL}/api/zones/records/add", params={
"token": token, "zone": zone, "domain": fqdn,
"type": "A", "ipAddress": ip, "ttl": "300", "overwrite": "true",
}).raise_for_status()Then the deletion pass: list the A records that exist, drop the ones NetBox no
longer knows about.
This only runs inside zones the flow is allowed to touch, set by
u/admin/dns_allowed_zones:
- left empty: every non-internal zone is managed;
- set: the sync never wanders outside that list;
- either way, Technitium’s own internal zones (
localhost,in-addr.arpa) are excluded automatically.
DNS conflicts: what is handled, what is not
Two situations look like a conflict. Only one is.
Several dns_name values on the same IP: no problem. desired is keyed by
FQDN, so web01.lab.example and www.lab.example both pointing at
10.60.20.10 are two distinct keys and two distinct A records. The deletion
pass also works by name, so removing one leaves the other alone. That is the
nominal case, not something being tolerated.
Two IPs carrying the same dns_name: one of them silently vanishes. Same
dictionary, seen from the other side:
desired[fqdn] = ip["address"].split("/")[0]The last IP read overwrites the previous one before Technitium is contacted
at all. overwrite=true has nothing to do with it: the add call only ever
receives one value for that name, so it is not replacing a competing record, it
is applying the only one it was given. NetBox sorts addresses by address, so in
practice the highest one wins, but that is a consequence of read order, not a
resolution rule. Nothing fails and nothing is reported: dns_records counts the
dictionary keys, so the losing IP appears nowhere in the run summary.
And nothing watches for this today. No alert, no collision counter. A
duplicated dns_name entered by mistake in NetBox would go unnoticed, and the
only symptom would be a host that does not resolve, reported by a person rather
than by the system. This is a gap spotted afterwards, not a trade-off weighed at
design time. The place to fix it is the read step, which has everything it needs
to detect the collision and fail, or at least surface it. It does not.
Sync DHCP: skip the runs that change nothing
The DHCP webhook listens on ipam.ipaddress, because reservations live on a
custom field of the IP rather than in a model of their own. So it also fires
when someone edits dns_name, which has nothing to do with DHCP.
Step one compares the before and after snapshots and stops early:
def main(object_type: str = None, snapshots: dict = None):
if object_type == "ipam.ipaddress":
snapshots = snapshots or {}
before = (snapshots.get("prechange") or {}).get("custom_fields", {}).get("mac_address")
after = (snapshots.get("postchange") or {}).get("custom_fields", {}).get("mac_address")
if before == after:
return {"status": "skipped",
"reason": "ipaddress change irrelevant to DHCP"}
...Downstream steps carry skip_if: results.read_netbox.status == 'skipped'. A
dns_name edit no longer triggers a pointless config-set. On a schedule or a
manual run object_type is empty, and the full sync happens.
Sync DHCP: build the subnets
for prefix in prefixes:
net = ipaddress.ip_network(prefix["prefix"], strict=False)
pools = [
{"pool": f"{r['start_address'].split('/')[0]} - {r['end_address'].split('/')[0]}"}
for r in ranges
if ip_in_network(r["start_address"].split("/")[0], str(net))
]
subnets.append({
"id": prefix["id"], # NetBox id, stable across runs
"subnet": str(net),
"pools": pools,
"option-data": [
{"name": "routers", "data": str(net.network_address + 1)},
{"name": "domain-name-servers", "data": kea_dns_servers},
],
"reservations": [],
})Reusing the NetBox object id as the Kea subnet id is a small detail that pays off. The id is stable across runs, so a subnet keeps its identity even when the array is rebuilt from scratch.
Sync DHCP: apply atomically
r = requests.post(KEA_URL, json={"command": "config-get",
"service": ["dhcp4"], "arguments": {}})
dhcp = r.json()[0]["arguments"]["Dhcp4"]
dhcp["subnet4"] = subnets # the sync owns this array
r = requests.post(KEA_URL, json={"command": "config-set",
"service": ["dhcp4"],
"arguments": {"Dhcp4": dhcp}})
result = r.json()[0]
if result["result"] != 0:
raise RuntimeError(f"Kea config-set failed: {result['text']}")Read the running config, replace one array, write it back. Two HTTP calls whatever the number of subnets, applied without restarting the service.
Active leases survive the reload. That is what makes running this every hour safe.
Sync Proxmox SDN
VLANs become VNets, one per VLAN, named vl<vid>:
for name, want in desired.items():
if name not in existing:
pve("POST", "/cluster/sdn/vnets",
{"vnet": name, "zone": sdn_zone,
"tag": want["vid"], "alias": want["alias"]})
else:
cur = existing[name]
if cur.get("tag") != want["vid"] or (cur.get("alias") or "") != want["alias"]:
pve("PUT", f"/cluster/sdn/vnets/{name}",
{"tag": want["vid"], "alias": want["alias"]})
for name in existing:
if name not in desired:
pve("DELETE", f"/cluster/sdn/vnets/{name}")
if created + updated + removed > 0:
pve("PUT", "/cluster/sdn", {}) # without this, nothing is appliedProxmox also wants an API token rather than the root password. One more provisioning step, run exactly once:
docker compose run --rm proxmox-init
# creates root@pam!sync and the SDN zone, prints the token value
# copy it into .env as PROXMOX_TOKEN_VALUE, then:
docker compose run --rm windmill-setupThe token is stored as a Windmill secret in the PVEAPIToken=user!name=value
form, so no script ever holds a password.
Triggers: event-driven, with a safety net
setup.py registers three NetBox event rules, each pointed at a flow:
| Flow | NetBox object types | Schedule |
|---|---|---|
u/admin/sync_dns | ipam.ipaddress | 0 0 * * * * |
u/admin/sync_dhcp | ipam.prefix, ipam.iprange, ipam.ipaddress | 0 0 * * * * |
u/admin/sync_proxmox | ipam.vlan | 0 0 * * * * |
The setup code detects the NetBox major version. It registers plain webhooks on 3.x, webhooks plus event rules on 4.x, so the same stack works on both.
Two triggers, one destination:
fast path safety net
───────── ──────────
someone saves a form hourly schedule
webhook fires in seconds catches whatever was lost
│ │
└──────────────┐ ┌──────────────┘
▼ ▼
┌──────────────────────┐
│ the same reconcile │
├──────────────────────┤
│ desired ← NetBox │
│ actual ← the target │
│ apply the difference │
└──────────────────────┘Webhooks are the fast path. A DNS change is live seconds after somebody saves the form.
The hourly schedule is the honest path, because things go wrong:
- webhooks get lost,
- containers restart,
- somebody edits the database directly.
The scheduled run does not care what happened in between, because it never computes a diff from events. It reads the desired state, reads the actual state, and makes the second match the first.
Verifying it works
DNS, end to end:
TOKEN=$(curl -sf "http://localhost:5380/api/user/login?user=admin&pass=${TECHNITIUM_PASSWORD}" | jq -r .token)
curl -sf "http://localhost:5380/api/zones/records/get?token=$TOKEN&zone=lab.example&domain=lab.example&listZone=true" \
| jq '[.response.records[] | select(.type=="A") | {name, ip: .rData.ipAddress}]'
# resolution from inside ddns-net, where the resolver actually listens
docker run --rm --network ddi-stack_ddns-net nicolaka/netshoot \
dig @technitium web01.lab.example A +shortDHCP:
curl -s -X POST http://localhost:8080 \
-H "Content-Type: application/json" \
-d '{"command":"config-get","service":["dhcp4"],"arguments":{}}' \
| jq '[.[] | .arguments.Dhcp4.subnet4[]
| {subnet, pools: [.pools[].pool],
reservations: [.reservations[]["ip-address"]]}]'
curl -s -X POST http://localhost:8080 \
-H "Content-Type: application/json" \
-d '{"command":"lease4-get-all","service":["dhcp4"],"arguments":{"subnets":[1]}}' \
| jq '.[] | .arguments.leases[]'Forcing a sync without touching NetBox:
TOKEN=$(curl -s -X POST http://localhost:8300/api/auth/login \
-H "Content-Type: application/json" \
-d "{\"email\":\"${WINDMILL_USER}\",\"password\":\"${WINDMILL_PASSWORD}\"}")
curl -X POST "http://localhost:8300/api/w/ddi/jobs/run/f/u/admin/sync_dns" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" -d '{}'Cost
| Operation | HTTP requests | Atomic |
|---|---|---|
| Read NetBox | Paginated, 200 items per request | - |
| Upsert DNS | 1 per record (overwrite=true) | No |
| Delete DNS | 1 per record | No |
| Sync DHCP | 1 config-get + 1 config-set | Yes |
| Sync Proxmox SDN | 1 GET zones + 1 GET vnets + N writes | No |
Roughly 800 addresses across three zones sync in under a second. That is not a performance achievement, it is a consequence of the shape. Reading NetBox is paginated and cheap. DHCP collapses into two calls at any size. Only DNS scales linearly with the record count, and at this size linear is free.
The number that actually matters is different. Deleting a VM is now one action: remove the IP in NetBox. The record goes, the reservation goes, and nothing anywhere still claims that address.
Taking this to production?
Everything above is a proof of concept: one of each component, one host, a few hundred addresses. Five jobs open up once it has to survive a node reboot and hold tens of thousands of addresses across several sites.
Making every component highly available
Good news first: the orchestrator is not in the data path. If Windmill dies, DNS still resolves and DHCP still hands out leases from the last applied state. That is not an outage, it is a freeze.
Which reframes the job. The control plane needs to be recoverable. The data plane needs to be redundant. Different problems.
| Component | What HA looks like |
|---|---|
| PostgreSQL | Streaming replication with automatic failover (Patroni, or a Kubernetes operator) |
| Redis | Sentinel, or a managed equivalent, for both the task and caching databases |
| NetBox | Stateless web tier, N replicas behind a load balancer; media on shared or object storage |
| rqworker | Queue-backed, so scale it horizontally; it is where webhook delivery actually happens |
| Windmill | Server and workers split (DISABLE_SERVER, NUM_WORKERS), state in its own replicated PostgreSQL |
| Technitium | Hidden primary written by the sync, N secondaries serving clients over AXFR/IXFR |
| Kea | The high_availability hook in hot-standby or load-balancing mode, with lease sync between peers |
Two traps in that table.
The NetBox worker is not optional, and it is easy to forget when scaling. Outgoing webhooks queue through it. An undersized pool just looks like “the sync is lagging”, with nothing obviously broken.
HA Kea breaks the config-set approach. Push the whole subnet4 array to
one peer and the other is stale. Push to both and it is not atomic across the
pair. At that point config stops being a file you push and becomes a database
both servers read. The scaling section below arrives at the same conclusion from
a different direction.
Splitting by site, region and tier
NetBox already has the partition keys: Region, Site Group, Site, Tenant.
Use them instead of inventing a tagging scheme.
The read side is a query filter, which is the easy half:
prefixes = netbox_get_all(f"ipam/prefixes/?tag={dhcp_tag}&site={site}", auth)The hard half is everything around it. A flow scoped to a site needs:
- its own credentials,
- its own allow-list,
- its own targets.
A mistake in one region must not be able to reach another.
In Windmill that maps onto worker tags: pin a region’s job to workers running in that region. It also solves reachability, when the central worker has no route to a remote management network. Blast radius becomes a deployment property rather than a hope.
Environment tiering (lab, pre-production, production) is a different axis. One
NetBox with a Tenant per tier works and keeps a single source of truth. But
only if the sync credentials differ per tier: what stops a lab change reaching
production must be authorisation, not a filter in a Python file anyone can edit.
Central control plane, local data plane
“One place to declare, many places to serve” splits cleanly by protocol, because DNS and DHCP do not relate to the network the same way.
DNS federates natively. A hidden primary that only the sync writes, and a secondary at every site:
writes AXFR / IXFR + NOTIFY, TSIG-signed
sync ───────────▶ hidden primary ──┬──▶ secondary, site A ──▶ clients
(serves nobody) ├──▶ secondary, site B ──▶ clients
└──▶ secondary, site C ──▶ clients
no write API exposed at allThe remotes are read-only by construction, not by policy: they expose no write API for anyone to misuse. A WAN outage degrades to serving the last transferred zone, which is almost always the right failure mode.
DHCP does not federate, because it is link-local. Two honest options:
A. relay to the centre
clients ──▶ gateway (ip helper-address) ══ WAN ══▶ central Kea
WAN down ⇒ no new leases at that site
B. a Kea per site, configured from the centre
clients ──▶ local Kea ◀══ config pushed (or pulled from a shared DB)
WAN down ⇒ the site keeps servingOption A is one config to maintain. The cost is that existing clients only keep
working until their lease expires, which turns valid-lifetime into a
deliberate availability decision rather than a default.
Option B survives WAN loss, at the cost of N servers to keep in step.
The rule underneath both is worth carrying: the sync writes to exactly one place per service, and everything else is a replica. Two writers, and you have rebuilt the problem this whole project exists to remove.
Syncing the change, not the whole estate
This is what the POC most obviously trades away. Every run reads every IP, rebuilds the desired state, and diffs. At a few hundred addresses that costs under a second. At forty thousand, on every single edit, it does not.
today every edit ──▶ read ALL 40 000 IPs ──▶ diff ──▶ apply
wanted one edit ──▶ read the event payload ──▶ 1 write
hourly ──▶ read ONE zone/site ──▶ diff ──▶ applyThe irony: NetBox already sends what is needed, and the flow throws it away. The
event payload carries event, the object, and snapshots.prechange /
snapshots.postchange. Today the DHCP flow only uses that to decide whether to
skip. It could use it to decide what to write:
| Change | What the payload gives you | What to write |
|---|---|---|
IP created with dns_name | postchange | One record added |
dns_name edited | prechange + postchange | One delete, one add |
| IP deleted | prechange only | One record deleted |
| Unrelated field edited | Both, identical on the fields you care about | Nothing |
Deletion is the case that forces the design. postchange is null, so the old
value has to come from prechange. Any incremental sync that ignores prechange
snapshots silently leaks stale records.
Four things have to come with it:
- Keep a reconcile, but scope it. Per-object sync drifts the first time a
webhook is lost. Keep the full pass, make it partial: one zone or one site per
run, round-robin. Or filter on
?last_updated__gte=with a watermark in Windmill state, remembering that a timestamp filter cannot see deletions. - Coalesce the events. A bulk import of 5000 addresses should not be 5000 flow runs. Batch on a short window, deduplicate by object id.
- Lock per key. Many small concurrent runs race in ways one big serial run never did. Windmill’s concurrency limits, keyed per zone or per subnet, are the cheap fix. And re-read the object rather than trusting a payload that may already be stale.
- Change the target protocol, not just the code. This is the real ceiling. A
whole-array
config-setis O(estate) however clever the flow is. Move reservations into a MySQL or PostgreSQL host backend that Kea reads directly and the push disappears entirely. ISC also sellssubnet_cmdsandhost_cmdspremium hooks for per-object edits. For DNS the equivalent is RFC 2136 dynamic updates, per-record by definition.
At scale the answer is not a faster sync. It is a target that accepts a delta.
The pieces the POC does not have
Nothing original here, and that is the problem: three standard gaps it would be dishonest to leave for the reader to infer.
Backup. There is no backup strategy today.
- NetBox’s PostgreSQL is the only irreplaceable piece. It is the source of truth, everything else derives from it.
- The Technitium and Kea configurations do not need backing up: one sync rebuilds
them entirely. Windmill’s PostgreSQL sits in between,
setup.pyrecreates the flows and the schedules, only the run history is gone. - DHCP leases are the real loss. Kea’s memfile lives in the
kea-leasesvolume: lose the volume and the reservations come back on the next sync because they come from NetBox, but the dynamic leases from the pool are not rebuildable. Kea restarts from a pool it believes is free while clients still hold their addresses, until the estate renews.
Secrets. The .env sits in clear text on the host, is injected as
environment variables, and is relayed into Windmill by setup.py. The sensitive
values do land there as is_secret variables, but everything upstream is
readable by anyone with a shell on the machine. That is a deliberate trade for a
homelab POC, where the .env and the host share one surface. Moving to Docker
secrets, to SOPS for encrypting the file in Git, or to a Vault whose values are
read at boot is standard wiring on a Compose stack, not a redesign.
Observability. Finding out that a flow failed means opening the Windmill UI and looking at the run history. Nobody is notified. The hourly schedule limits the damage (the next run catches up on a one-off failure), but a lasting failure, an expired token or an unreachable target, stays invisible until it shows up somewhere else. Windmill’s error handler pointed at an alerting webhook, or a periodic poll of failed runs through its API, are the two obvious answers. Neither requires touching the flows.
What I would tell myself before starting
Decide what the automation may delete, on day one. A sync that only adds is
not a sync, it is an import. But one that deletes everything it does not
recognise will eventually eat a record you created by hand and forgot.
dns_allowed_zones exists because the first version had no such limit.
Make the setup step idempotent. Workspace, variables, flows, schedules, webhooks: all created by a script that can be replayed at will. Anything provisioned by hand becomes the one thing nobody can rebuild.
Keep the flow steps small and separate. Read, compute, apply. Three steps tell you where a run failed in one glance. One step makes you read logs.
Do not publish DNS and DHCP on the host. They are infrastructure for the stack, not services for whatever network the server running the POC happens to sit on. Keeping them unpublished was the easiest security decision in the whole project.
That decision belongs to the Docker Compose POC, but the principle transfers unchanged to production: isolate the network services in a dedicated VLAN and expose only what has to be exposed. In practice that means a DHCP relay rather than a server per VLAN, DoH (DNS over HTTPS) for client resolution, and small read-only DNS servers per zone fed by AXFR. The goal is the same either way: standard behaviour, less network configuration, and certainly not every port open on every VLAN.
The pattern does not stop at DNS and DHCP: each new target is one more read, compute, apply flow. vCenter port groups, switch configurations NetBox already knows how to render from its config templates, firewall address objects (the objects, not the rules, which have to stay under human review). I have implemented none of them, so I will say nothing more about them.
One thing is worth anticipating before getting there: the more consumers there
are, the more NetBox data quality becomes critical. With one, a malformed
dns_name is one bad record. With six, it propagates everywhere at once.

