← All notes
A row of datacenter racks lit in blue, each carrying a glowing panel with the NetBox, Technitium, Kea and Proxmox logos
NoteBy Kilian Toulliou27 min read

One source of truth for DNS, DHCP and IPAM

A reproducible self-hosted DDI proof of concept where NetBox drives Technitium DNS, Kea DHCP and the Proxmox VE SDN through event-driven Windmill flows. Full Docker Compose stack and configuration included.

Every lab I run catches the same disease. One IP address ends up living in four places: a DNS record, a DHCP reservation, a VLAN on the Proxmox host, and a spreadsheet that stopped being true months ago.

Nothing breaks until you delete a machine. Then three of those four keep pointing at it.

The fix I settled on: one source of truth, everything else generated from it. All self-hosted, nothing leaving the rack. What follows is enough to rebuild the whole thing from an empty directory.

Architecture

   source of truth              orchestrator            enforced state
 ┌────────────────────┐     ┌──────────────────┐     ┌────────────────────┐
 │      NetBox        │     │    Windmill      │ ──▶ │  Technitium DNS    │
 │                    │     │                  │     ├────────────────────┤
 │ prefixes [dhcp]    │ ──▶ │ flow: sync DNS   │ ──▶ │  Kea DHCP4         │
 │ IPs (dns_name)     │     │ flow: sync DHCP  │     ├────────────────────┤
 │ IPs (mac_address)  │     │ flow: sync SDN   │ ──▶ │  Proxmox VE SDN    │
 │ VLANs              │     │                  │     │  (VNets per VLAN)  │
 └────────────────────┘     └──────────────────┘     └────────────────────┘
        ▲                     ▲            ▲
        │                     │            │
     humans            webhook on change   hourly reconcile
RoleWhat runs itWhy this one
IPAMNetBoxReal data model, proper API, changelog on every object
OrchestrationWindmillFlows as versioned Python steps, plus a UI to run them
DNSTechnitiumFull HTTP API, sane zone handling, lightweight
DHCPKeaConfig is JSON and reloadable over a control socket
Network overlayProxmox VE SDNVLANs already exist in NetBox, so they may as well match

Three Docker networks, and the split matters:

  • netbox-net: NetBox, its Postgres and its Redis.
  • ddns-net: DNS, DHCP and Proxmox.
  • windmill-net: Windmill and its own database.

Windmill is the only container on all three. It is therefore the only thing that can reach both the source of truth and the targets.

Ports 53 and 67 are not published on the host, deliberately. The resolver and the lease server answer inside ddns-net and nowhere else. That one decision kills an entire class of “why is the Docker host resolving through the lab” problems.

Why Windmill, after two attempts that were not?

Windmill is the third orchestrator this POC ran on. The first two were not mistakes. They were the shortest path to understanding what the job needed.

First attempt: a Python container on cron

A sync/ directory, three scripts, and an entrypoint that looped over them:

python
SYNC_INTERVAL = int(os.environ.get("SYNC_INTERVAL", "60"))

def job() -> None:
    run_sync("sync_dns")
    run_sync("sync_dhcp")
    run_sync("sync_proxmox")

job()
schedule.every(SYNC_INTERVAL).seconds.do(job)

while True:
    schedule.run_pending()
    time.sleep(1)

One SYNC_INTERVAL in .env, 60 seconds by default. It worked. The sync logic inside it is essentially what still runs today.

What it could not do was tell me anything:

  • Every run is a full poll. Nothing changed for six hours? It still hit every endpoint 360 times. And one interval had to serve three very different jobs.
  • Failures scroll past. run_sync catches, logs and carries on. That part is right: one broken target must not stop the other two. But the only trace is a line in docker logs, gone at the next restart. “Did the DHCP sync succeed at 14:03, and what did it push?” has no answer.

The rest follows: nothing to run on demand, and read, compute and apply are one function call. Debugging meant adding a print and waiting another 60 seconds.

Second attempt: n8n

Next I moved the same logic into n8n, driven by NetBox webhooks. The workflows went into the repository as n8n/workflows/*.json.

Event-driven was the right call, and it stayed. n8n as infrastructure as code was not.

Look at what a workflow file actually is. Here is one node, as committed:

json
{
  "id": "22222222-2222-4222-8222-222222222222",
  "name": "Lire Netbox",
  "type": "n8n-nodes-base.code",
  "typeVersion": 2,
  "position": [550, 300],
  "parameters": {
    "mode": "runOnceForAllItems",
    "jsCode": "const NETBOX_URL = 'http://netbox:8080';\nconst NETBOX_AUTH = '__NETBOX_AUTH__';\n\nasync function netboxGetAll(path) {\n  const items = [];\n  let url = `${NETBOX_URL}/api/${path}...
  }
}

Every real line of code sits inside one JSON string, newlines escaped. That single detail poisons everything:

  • It cannot be reviewed. Change one character and git diff shows one enormous modified line, and no tooling reaches what is inside it: no highlighting, no linter, no types. The editor sees a string.
  • Version control holds UI state. "position": [550, 300] is where the box sits on the canvas. Move a box in the browser, get a diff.

Editing in the browser is pleasant. Right up to the moment two sources of truth exist, and you are diffing generated JSON by hand.

What Windmill changed

Windmill keeps what n8n got right: event-driven, a DAG (directed acyclic graph) you can look at, per-step results in a UI. It drops the rest, because a flow step is a normal Python file on disk:

windmill/scripts/sync_dhcp/
├── read_netbox.py        # read
├── compute_config.py     # compute
└── apply_kea.py          # apply

setup.py reads those files and pushes them into Windmill.

ConcernCron containern8nWindmill
TriggerFixed interval onlyWebhook + schedule + manualWebhook + schedule + manual
Logic in GitReal .py filesEscaped JSON stringsReal .py files
Diff you can reviewYesNoYes
Per-step resultsNoYesYes
Run historydocker logsYesYes
SecretsEnv varsPlaceholder substitutionTyped secret variables
RBACNoneEnterprise editionUsers, groups, folders
Maintenance burdenOne image, no stateA database, plus workflows to re-exportA dedicated PostgreSQL

What that buys:

  • The code is linted and diffed like any other Python.
  • Each step’s input, output and duration survives the run.
  • A flow can be replayed from the UI while debugging.
  • Secrets are variables, not strings interpolated into source at boot.

The honest caveat: this is not free. Windmill wants its own PostgreSQL, which is one more database to run and back up. If your sync is a single script with no branching and nobody else ever reads it, the cron container was fine. I would not talk you out of it.

Repository layout

.
├── compose.yaml
├── .env
├── config/
│   ├── kea/
│   │   ├── kea-dhcp4.conf          # subnet4 is empty, the sync fills it
│   │   └── kea-ctrl-agent.conf     # Unix socket exposed over HTTP
│   ├── netbox/
│   │   └── configuration.py
│   └── proxmox/
│       └── init_sdn.py             # one-shot: API token + SDN zone
└── windmill/
    ├── setup.py                    # deploys variables, flows, schedules, webhooks
    └── scripts/
        ├── sync_dns/{read_netbox,sync_technitium}.py
        ├── sync_dhcp/{read_netbox,compute_config,apply_kea}.py
        └── sync_proxmox/{read_netbox_vlans,sync_proxmox_sdn}.py

The Docker Compose stack

The whole POC lives in one compose.yaml. It is long, so here are the parts that carry a decision. NetBox and its dependencies first:

yaml
name: ddi-stack

networks:
  netbox-net:
  ddns-net:
  windmill-net:

services:
  postgres:
    image: postgres:16-alpine
    restart: unless-stopped
    networks: [netbox-net]
    environment:
      POSTGRES_DB: netbox
      POSTGRES_USER: netbox
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
    volumes:
      - postgres-data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U netbox"]
      interval: 10s
      retries: 5

  netbox:
    image: netboxcommunity/netbox:latest
    restart: unless-stopped
    networks: [netbox-net, ddns-net]
    ports:
      - "8000:8080"
    environment:
      SECRET_KEY: ${NETBOX_SECRET_KEY}
      NETBOX_TOKEN_PEPPER: ${NETBOX_TOKEN_PEPPER}
      DB_HOST: postgres
      REDIS_HOST: redis
      REDIS_CACHE_HOST: redis
      REDIS_CACHE_DATABASE: "1"
    volumes:
      - ./config/netbox/configuration.py:/etc/netbox/config/configuration.py:ro
    depends_on:
      postgres: { condition: service_healthy }
      redis: { condition: service_healthy }
    healthcheck:
      test: ["CMD-SHELL", "curl -sf http://localhost:8080/login/ || exit 1"]
      interval: 30s
      retries: 10
      start_period: 120s

Two things worth copying.

start_period: 120s on the NetBox healthcheck. First boot runs migrations and takes minutes. Without it, Compose calls the container unhealthy and everything downstream gives up.

DNS and DHCP:

yaml
  technitium:
    image: technitium/dns-server:latest
    networks: [ddns-net]
    ports:
      - "5380:5380"          # web UI and API only, never 53
    environment:
      DNS_SERVER_ADMIN_PASSWORD: ${TECHNITIUM_PASSWORD}
    volumes:
      - technitium-data:/etc/dns

  kea:
    image: jonasal/kea-dhcp4:2
    command: -c /etc/kea/kea-dhcp4.conf
    networks: [ddns-net]
    volumes:
      - kea-leases:/kea/leases
      - kea-sockets:/kea/sockets      # shared with the control agent
      - ./config/kea:/etc/kea
    healthcheck:
      test: ["CMD-SHELL", "test -S /kea/sockets/kea-dhcp4-ctrl.sock || exit 1"]
      interval: 15s
      start_period: 20s

  kea-ctrl-agent:
    image: jonasal/kea-ctrl-agent:2
    command: -c /etc/kea/kea-ctrl-agent.conf
    networks: [ddns-net]
    ports:
      - "8080:8000"
    volumes:
      - kea-sockets:/kea/sockets
      - ./config/kea:/etc/kea
    depends_on:
      kea: { condition: service_healthy }

Kea is split into two containers deliberately:

  ┌──────────────┐   Unix socket    ┌────────────────┐   HTTP :8000
  │  kea-dhcp4   │◀────────────────▶│ kea-ctrl-agent │◀──────────────  the sync
  │  speaks DHCP │  kea-sockets vol │  speaks HTTP   │
  └──────────────┘                  └────────────────┘
  • kea-dhcp4 speaks DHCP and exposes a Unix control socket.
  • kea-ctrl-agent is the only thing that turns that socket into HTTP.
  • They share it through the kea-sockets volume.
  • The healthcheck on the socket file stops the agent starting before there is anything to talk to.

Windmill and its setup job:

yaml
  windmill:
    image: ghcr.io/windmill-labs/windmill:main
    networks: [windmill-net, netbox-net, ddns-net]   # the only bridge
    ports:
      - "8300:8000"
    environment:
      DATABASE_URL: postgresql://windmill:${WINDMILL_DB_PASSWORD}@windmill-db/windmill
      BASE_INTERNAL_URL: http://windmill:8000
      SUPERADMIN_SECRET: ${WINDMILL_SUPERADMIN_SECRET}
      NUM_WORKERS: "1"
    depends_on:
      windmill-db: { condition: service_healthy }

  windmill-setup:
    image: python:3.12-slim
    restart: on-failure                 # retries until Windmill answers
    command: sh -c "pip install -q requests && python /setup.py"
    networks: [windmill-net, netbox-net, ddns-net]
    volumes:
      - ./windmill/scripts:/windmill-scripts:ro
      - ./windmill/setup.py:/setup.py:ro
    environment:
      WINDMILL_URL: http://windmill:8000
      NETBOX_URL: http://netbox:8080
      NETBOX_TOKEN: ${NETBOX_TOKEN}
      DNS_ALLOWED_ZONES: ${DNS_ALLOWED_ZONES:-}
      KEA_DNS_SERVERS: ${KEA_DNS_SERVERS:-9.9.9.9, 149.112.112.112}
      DHCP_TAG: ${DHCP_TAG:-dhcp}
    depends_on:
      windmill: { condition: service_healthy }
      netbox: { condition: service_healthy }

restart: on-failure on the setup container is the cheap way to handle ordering. It exits non-zero while Windmill boots, Compose restarts it, it eventually succeeds. No wait loop to write.

Kea, configured to be overwritten

The Kea config on disk is deliberately almost empty:

json
{
  "Dhcp4": {
    "interfaces-config": {
      "interfaces": ["*"],
      "dhcp-socket-type": "udp"
    },
    "control-socket": {
      "socket-type": "unix",
      "socket-name": "/kea/sockets/kea-dhcp4-ctrl.sock"
    },
    "lease-database": {
      "type": "memfile",
      "persist": true,
      "name": "/kea/leases/dhcp4.leases"
    },
    "hooks-libraries": [
      { "library": "/usr/local/lib/kea/hooks/libdhcp_lease_cmds.so" }
    ],
    "valid-lifetime": 4000,
    "subnet4": []
  }
}

"subnet4": [] is the point. No subnet is ever written by hand. The sync owns that array outright, and that is what makes “delete the prefix in NetBox” actually remove the subnet.

libdhcp_lease_cmds.so is loaded so lease commands work over the control channel. The agent just bridges the socket:

json
{
  "Control-agent": {
    "http-host": "0.0.0.0",
    "http-port": 8000,
    "control-sockets": {
      "dhcp4": {
        "socket-type": "unix",
        "socket-name": "/kea/sockets/kea-dhcp4-ctrl.sock"
      }
    }
  }
}

Bringing it up

bash
cp .env.example .env
# fill in NETBOX_SECRET_KEY, NETBOX_TOKEN_PEPPER, POSTGRES_PASSWORD,
# TECHNITIUM_PASSWORD, WINDMILL_SUPERADMIN_SECRET, WINDMILL_PASSWORD

docker compose up -d
docker compose logs -f netbox      # migrations, 2 to 3 minutes on first boot

Create the NetBox superuser and an API token:

bash
docker compose exec -e DJANGO_SUPERUSER_PASSWORD=change-me netbox \
  /opt/netbox/venv/bin/python /opt/netbox/netbox/manage.py \
  createsuperuser --username admin --email admin@lab.example --noinput

docker compose exec netbox /opt/netbox/venv/bin/python \
  /opt/netbox/netbox/manage.py shell -c "
from users.models import Token
from django.contrib.auth import get_user_model
u = get_user_model().objects.get(username='admin')
print('NETBOX_TOKEN=' + Token.objects.create(user=u).key)
"

Put that token in .env as NETBOX_TOKEN, then replay the setup job:

bash
docker compose run --rm windmill-setup

That one command does everything:

  • creates the Windmill workspace,
  • pushes the variables, secret and plain,
  • deploys the three flows from windmill/scripts/,
  • registers the hourly schedules,
  • creates the NetBox event rules.

Nothing is clicked in a UI. It is idempotent, so re-running it updates instead of duplicating.

Declaring the data in NetBox

Four fields carry everything.

Create the dhcp tag once (Customization → Tags), and the mac_address custom field on IPAM > IP address (Customization → Custom Fields, type Text). Then:

NetBox objectValueTagBecomes
Prefix10.60.20.0/24dhcpKea subnet, router option .1
IP range.100 to .200dhcpDynamic pool inside that subnet
IP address10.60.20.10-A record when dns_name is set
IP address+ mac_address-Fixed reservation in Kea

Creating a host that gets both a record and a reservation:

bash
curl -s -X POST http://localhost:8000/api/ipam/ip-addresses/ \
  -H "Authorization: Token ${NETBOX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
        "address": "10.60.20.10/24",
        "status": "active",
        "dns_name": "web01.lab.example",
        "custom_fields": {"mac_address": "aa:bb:cc:00:11:22"}
      }'

Nobody fills in the default gateway. The flow computes it as the first host of the network. One less value to get wrong.

The flows, step by step

This is the heart of the project, and it is humbler than it sounds: make a few APIs talk to each other, on a trigger or on a schedule. No wheel is reinvented here. Data is translated from one API into another, and the translation is kept true.

Every flow is the same three-step DAG. That shape is the whole debugging story:

   trigger  (webhook │ schedule │ manual run)
      │
      ▼
 ┌──────────────┐
 │ read NetBox  │   failed here? the source gave me nothing usable
 └──────┬───────┘
        ▼
 ┌──────────────┐
 │ compute      │   failed here? I built the wrong config
 └──────┬───────┘
        ▼
 ┌──────────────┐
 │ apply        │   failed here? the target rejected it
 └──────────────┘

One glance at which box is red tells you where to look. That is worth more than it sounds.

Sync DNS: read NetBox

Step one is a paginated read, nothing else:

python
NETBOX_URL = "http://netbox:8080"

def netbox_get_all(path: str, auth: str) -> list:
    items, sep = [], "&" if "?" in path else "?"
    url = f"{NETBOX_URL}/api/{path}{sep}limit=200"
    while url:
        r = requests.get(url, headers={"Authorization": auth})
        r.raise_for_status()
        data = r.json()
        items.extend(data.get("results", []))
        url = data.get("next")
    return items

def main():
    auth = f"Token {wmill.get_variable('u/admin/netbox_token')}"
    desired = {}
    for ip in netbox_get_all("ipam/ip-addresses/", auth):
        if not ip.get("dns_name"):
            continue
        desired[ip["dns_name"].strip().rstrip(".")] = ip["address"].split("/")[0]
    return {"desired": desired, "dns_records": len(desired)}

Sync DNS: apply to Technitium

Step two picks the zone by longest suffix match. So web01.dev.lab.example lands in dev.lab.example when that zone exists, and in lab.example otherwise:

python
def find_zone(fqdn: str, zones: set) -> str | None:
    best = None
    for z in zones:
        if fqdn == z or fqdn.endswith(f".{z}"):
            if best is None or len(z) > len(best):
                best = z
    return best

The write is an upsert, so replaying the sync is a no-op:

python
requests.get(f"{TECHNITIUM_URL}/api/zones/records/add", params={
    "token": token, "zone": zone, "domain": fqdn,
    "type": "A", "ipAddress": ip, "ttl": "300", "overwrite": "true",
}).raise_for_status()

Then the deletion pass: list the A records that exist, drop the ones NetBox no longer knows about.

This only runs inside zones the flow is allowed to touch, set by u/admin/dns_allowed_zones:

  • left empty: every non-internal zone is managed;
  • set: the sync never wanders outside that list;
  • either way, Technitium’s own internal zones (localhost, in-addr.arpa) are excluded automatically.

DNS conflicts: what is handled, what is not

Two situations look like a conflict. Only one is.

Several dns_name values on the same IP: no problem. desired is keyed by FQDN, so web01.lab.example and www.lab.example both pointing at 10.60.20.10 are two distinct keys and two distinct A records. The deletion pass also works by name, so removing one leaves the other alone. That is the nominal case, not something being tolerated.

Two IPs carrying the same dns_name: one of them silently vanishes. Same dictionary, seen from the other side:

python
desired[fqdn] = ip["address"].split("/")[0]

The last IP read overwrites the previous one before Technitium is contacted at all. overwrite=true has nothing to do with it: the add call only ever receives one value for that name, so it is not replacing a competing record, it is applying the only one it was given. NetBox sorts addresses by address, so in practice the highest one wins, but that is a consequence of read order, not a resolution rule. Nothing fails and nothing is reported: dns_records counts the dictionary keys, so the losing IP appears nowhere in the run summary.

And nothing watches for this today. No alert, no collision counter. A duplicated dns_name entered by mistake in NetBox would go unnoticed, and the only symptom would be a host that does not resolve, reported by a person rather than by the system. This is a gap spotted afterwards, not a trade-off weighed at design time. The place to fix it is the read step, which has everything it needs to detect the collision and fail, or at least surface it. It does not.

Sync DHCP: skip the runs that change nothing

The DHCP webhook listens on ipam.ipaddress, because reservations live on a custom field of the IP rather than in a model of their own. So it also fires when someone edits dns_name, which has nothing to do with DHCP.

Step one compares the before and after snapshots and stops early:

python
def main(object_type: str = None, snapshots: dict = None):
    if object_type == "ipam.ipaddress":
        snapshots = snapshots or {}
        before = (snapshots.get("prechange") or {}).get("custom_fields", {}).get("mac_address")
        after = (snapshots.get("postchange") or {}).get("custom_fields", {}).get("mac_address")
        if before == after:
            return {"status": "skipped",
                    "reason": "ipaddress change irrelevant to DHCP"}
    ...

Downstream steps carry skip_if: results.read_netbox.status == 'skipped'. A dns_name edit no longer triggers a pointless config-set. On a schedule or a manual run object_type is empty, and the full sync happens.

Sync DHCP: build the subnets

python
for prefix in prefixes:
    net = ipaddress.ip_network(prefix["prefix"], strict=False)
    pools = [
        {"pool": f"{r['start_address'].split('/')[0]} - {r['end_address'].split('/')[0]}"}
        for r in ranges
        if ip_in_network(r["start_address"].split("/")[0], str(net))
    ]
    subnets.append({
        "id": prefix["id"],                       # NetBox id, stable across runs
        "subnet": str(net),
        "pools": pools,
        "option-data": [
            {"name": "routers", "data": str(net.network_address + 1)},
            {"name": "domain-name-servers", "data": kea_dns_servers},
        ],
        "reservations": [],
    })

Reusing the NetBox object id as the Kea subnet id is a small detail that pays off. The id is stable across runs, so a subnet keeps its identity even when the array is rebuilt from scratch.

Sync DHCP: apply atomically

python
r = requests.post(KEA_URL, json={"command": "config-get",
                                 "service": ["dhcp4"], "arguments": {}})
dhcp = r.json()[0]["arguments"]["Dhcp4"]
dhcp["subnet4"] = subnets                      # the sync owns this array

r = requests.post(KEA_URL, json={"command": "config-set",
                                 "service": ["dhcp4"],
                                 "arguments": {"Dhcp4": dhcp}})
result = r.json()[0]
if result["result"] != 0:
    raise RuntimeError(f"Kea config-set failed: {result['text']}")

Read the running config, replace one array, write it back. Two HTTP calls whatever the number of subnets, applied without restarting the service.

Active leases survive the reload. That is what makes running this every hour safe.

Sync Proxmox SDN

VLANs become VNets, one per VLAN, named vl<vid>:

python
for name, want in desired.items():
    if name not in existing:
        pve("POST", "/cluster/sdn/vnets",
            {"vnet": name, "zone": sdn_zone,
             "tag": want["vid"], "alias": want["alias"]})
    else:
        cur = existing[name]
        if cur.get("tag") != want["vid"] or (cur.get("alias") or "") != want["alias"]:
            pve("PUT", f"/cluster/sdn/vnets/{name}",
                {"tag": want["vid"], "alias": want["alias"]})

for name in existing:
    if name not in desired:
        pve("DELETE", f"/cluster/sdn/vnets/{name}")

if created + updated + removed > 0:
    pve("PUT", "/cluster/sdn", {})     # without this, nothing is applied

Proxmox also wants an API token rather than the root password. One more provisioning step, run exactly once:

bash
docker compose run --rm proxmox-init
# creates root@pam!sync and the SDN zone, prints the token value
# copy it into .env as PROXMOX_TOKEN_VALUE, then:
docker compose run --rm windmill-setup

The token is stored as a Windmill secret in the PVEAPIToken=user!name=value form, so no script ever holds a password.

Triggers: event-driven, with a safety net

setup.py registers three NetBox event rules, each pointed at a flow:

FlowNetBox object typesSchedule
u/admin/sync_dnsipam.ipaddress0 0 * * * *
u/admin/sync_dhcpipam.prefix, ipam.iprange, ipam.ipaddress0 0 * * * *
u/admin/sync_proxmoxipam.vlan0 0 * * * *

The setup code detects the NetBox major version. It registers plain webhooks on 3.x, webhooks plus event rules on 4.x, so the same stack works on both.

Two triggers, one destination:

  fast path                              safety net
  ─────────                              ──────────
  someone saves a form                   hourly schedule
  webhook fires in seconds               catches whatever was lost
          │                                      │
          └──────────────┐        ┌──────────────┘
                         ▼        ▼
                 ┌──────────────────────┐
                 │  the same reconcile  │
                 ├──────────────────────┤
                 │ desired ← NetBox     │
                 │ actual  ← the target │
                 │ apply the difference │
                 └──────────────────────┘

Webhooks are the fast path. A DNS change is live seconds after somebody saves the form.

The hourly schedule is the honest path, because things go wrong:

  • webhooks get lost,
  • containers restart,
  • somebody edits the database directly.

The scheduled run does not care what happened in between, because it never computes a diff from events. It reads the desired state, reads the actual state, and makes the second match the first.

Verifying it works

DNS, end to end:

bash
TOKEN=$(curl -sf "http://localhost:5380/api/user/login?user=admin&pass=${TECHNITIUM_PASSWORD}" | jq -r .token)

curl -sf "http://localhost:5380/api/zones/records/get?token=$TOKEN&zone=lab.example&domain=lab.example&listZone=true" \
  | jq '[.response.records[] | select(.type=="A") | {name, ip: .rData.ipAddress}]'

# resolution from inside ddns-net, where the resolver actually listens
docker run --rm --network ddi-stack_ddns-net nicolaka/netshoot \
  dig @technitium web01.lab.example A +short

DHCP:

bash
curl -s -X POST http://localhost:8080 \
  -H "Content-Type: application/json" \
  -d '{"command":"config-get","service":["dhcp4"],"arguments":{}}' \
  | jq '[.[] | .arguments.Dhcp4.subnet4[]
         | {subnet, pools: [.pools[].pool],
            reservations: [.reservations[]["ip-address"]]}]'

curl -s -X POST http://localhost:8080 \
  -H "Content-Type: application/json" \
  -d '{"command":"lease4-get-all","service":["dhcp4"],"arguments":{"subnets":[1]}}' \
  | jq '.[] | .arguments.leases[]'

Forcing a sync without touching NetBox:

bash
TOKEN=$(curl -s -X POST http://localhost:8300/api/auth/login \
  -H "Content-Type: application/json" \
  -d "{\"email\":\"${WINDMILL_USER}\",\"password\":\"${WINDMILL_PASSWORD}\"}")

curl -X POST "http://localhost:8300/api/w/ddi/jobs/run/f/u/admin/sync_dns" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" -d '{}'

Cost

OperationHTTP requestsAtomic
Read NetBoxPaginated, 200 items per request-
Upsert DNS1 per record (overwrite=true)No
Delete DNS1 per recordNo
Sync DHCP1 config-get + 1 config-setYes
Sync Proxmox SDN1 GET zones + 1 GET vnets + N writesNo

Roughly 800 addresses across three zones sync in under a second. That is not a performance achievement, it is a consequence of the shape. Reading NetBox is paginated and cheap. DHCP collapses into two calls at any size. Only DNS scales linearly with the record count, and at this size linear is free.

The number that actually matters is different. Deleting a VM is now one action: remove the IP in NetBox. The record goes, the reservation goes, and nothing anywhere still claims that address.

Taking this to production?

Everything above is a proof of concept: one of each component, one host, a few hundred addresses. Five jobs open up once it has to survive a node reboot and hold tens of thousands of addresses across several sites.

Making every component highly available

Good news first: the orchestrator is not in the data path. If Windmill dies, DNS still resolves and DHCP still hands out leases from the last applied state. That is not an outage, it is a freeze.

Which reframes the job. The control plane needs to be recoverable. The data plane needs to be redundant. Different problems.

ComponentWhat HA looks like
PostgreSQLStreaming replication with automatic failover (Patroni, or a Kubernetes operator)
RedisSentinel, or a managed equivalent, for both the task and caching databases
NetBoxStateless web tier, N replicas behind a load balancer; media on shared or object storage
rqworkerQueue-backed, so scale it horizontally; it is where webhook delivery actually happens
WindmillServer and workers split (DISABLE_SERVER, NUM_WORKERS), state in its own replicated PostgreSQL
TechnitiumHidden primary written by the sync, N secondaries serving clients over AXFR/IXFR
KeaThe high_availability hook in hot-standby or load-balancing mode, with lease sync between peers

Two traps in that table.

The NetBox worker is not optional, and it is easy to forget when scaling. Outgoing webhooks queue through it. An undersized pool just looks like “the sync is lagging”, with nothing obviously broken.

HA Kea breaks the config-set approach. Push the whole subnet4 array to one peer and the other is stale. Push to both and it is not atomic across the pair. At that point config stops being a file you push and becomes a database both servers read. The scaling section below arrives at the same conclusion from a different direction.

Splitting by site, region and tier

NetBox already has the partition keys: Region, Site Group, Site, Tenant. Use them instead of inventing a tagging scheme.

The read side is a query filter, which is the easy half:

python
prefixes = netbox_get_all(f"ipam/prefixes/?tag={dhcp_tag}&site={site}", auth)

The hard half is everything around it. A flow scoped to a site needs:

  • its own credentials,
  • its own allow-list,
  • its own targets.

A mistake in one region must not be able to reach another.

In Windmill that maps onto worker tags: pin a region’s job to workers running in that region. It also solves reachability, when the central worker has no route to a remote management network. Blast radius becomes a deployment property rather than a hope.

Environment tiering (lab, pre-production, production) is a different axis. One NetBox with a Tenant per tier works and keeps a single source of truth. But only if the sync credentials differ per tier: what stops a lab change reaching production must be authorisation, not a filter in a Python file anyone can edit.

Central control plane, local data plane

“One place to declare, many places to serve” splits cleanly by protocol, because DNS and DHCP do not relate to the network the same way.

DNS federates natively. A hidden primary that only the sync writes, and a secondary at every site:

          writes            AXFR / IXFR + NOTIFY, TSIG-signed
  sync ───────────▶ hidden primary ──┬──▶ secondary, site A ──▶ clients
                    (serves nobody)  ├──▶ secondary, site B ──▶ clients
                                     └──▶ secondary, site C ──▶ clients
                                          no write API exposed at all

The remotes are read-only by construction, not by policy: they expose no write API for anyone to misuse. A WAN outage degrades to serving the last transferred zone, which is almost always the right failure mode.

DHCP does not federate, because it is link-local. Two honest options:

 A. relay to the centre
    clients ──▶ gateway (ip helper-address) ══ WAN ══▶ central Kea
    WAN down ⇒ no new leases at that site

 B. a Kea per site, configured from the centre
    clients ──▶ local Kea ◀══ config pushed (or pulled from a shared DB)
    WAN down ⇒ the site keeps serving

Option A is one config to maintain. The cost is that existing clients only keep working until their lease expires, which turns valid-lifetime into a deliberate availability decision rather than a default.

Option B survives WAN loss, at the cost of N servers to keep in step.

The rule underneath both is worth carrying: the sync writes to exactly one place per service, and everything else is a replica. Two writers, and you have rebuilt the problem this whole project exists to remove.

Syncing the change, not the whole estate

This is what the POC most obviously trades away. Every run reads every IP, rebuilds the desired state, and diffs. At a few hundred addresses that costs under a second. At forty thousand, on every single edit, it does not.

 today        every edit ──▶ read ALL 40 000 IPs ──▶ diff ──▶ apply
 wanted       one edit   ──▶ read the event payload ──▶ 1 write
              hourly     ──▶ read ONE zone/site    ──▶ diff ──▶ apply

The irony: NetBox already sends what is needed, and the flow throws it away. The event payload carries event, the object, and snapshots.prechange / snapshots.postchange. Today the DHCP flow only uses that to decide whether to skip. It could use it to decide what to write:

ChangeWhat the payload gives youWhat to write
IP created with dns_namepostchangeOne record added
dns_name editedprechange + postchangeOne delete, one add
IP deletedprechange onlyOne record deleted
Unrelated field editedBoth, identical on the fields you care aboutNothing

Deletion is the case that forces the design. postchange is null, so the old value has to come from prechange. Any incremental sync that ignores prechange snapshots silently leaks stale records.

Four things have to come with it:

  • Keep a reconcile, but scope it. Per-object sync drifts the first time a webhook is lost. Keep the full pass, make it partial: one zone or one site per run, round-robin. Or filter on ?last_updated__gte= with a watermark in Windmill state, remembering that a timestamp filter cannot see deletions.
  • Coalesce the events. A bulk import of 5000 addresses should not be 5000 flow runs. Batch on a short window, deduplicate by object id.
  • Lock per key. Many small concurrent runs race in ways one big serial run never did. Windmill’s concurrency limits, keyed per zone or per subnet, are the cheap fix. And re-read the object rather than trusting a payload that may already be stale.
  • Change the target protocol, not just the code. This is the real ceiling. A whole-array config-set is O(estate) however clever the flow is. Move reservations into a MySQL or PostgreSQL host backend that Kea reads directly and the push disappears entirely. ISC also sells subnet_cmds and host_cmds premium hooks for per-object edits. For DNS the equivalent is RFC 2136 dynamic updates, per-record by definition.

At scale the answer is not a faster sync. It is a target that accepts a delta.

The pieces the POC does not have

Nothing original here, and that is the problem: three standard gaps it would be dishonest to leave for the reader to infer.

Backup. There is no backup strategy today.

  • NetBox’s PostgreSQL is the only irreplaceable piece. It is the source of truth, everything else derives from it.
  • The Technitium and Kea configurations do not need backing up: one sync rebuilds them entirely. Windmill’s PostgreSQL sits in between, setup.py recreates the flows and the schedules, only the run history is gone.
  • DHCP leases are the real loss. Kea’s memfile lives in the kea-leases volume: lose the volume and the reservations come back on the next sync because they come from NetBox, but the dynamic leases from the pool are not rebuildable. Kea restarts from a pool it believes is free while clients still hold their addresses, until the estate renews.

Secrets. The .env sits in clear text on the host, is injected as environment variables, and is relayed into Windmill by setup.py. The sensitive values do land there as is_secret variables, but everything upstream is readable by anyone with a shell on the machine. That is a deliberate trade for a homelab POC, where the .env and the host share one surface. Moving to Docker secrets, to SOPS for encrypting the file in Git, or to a Vault whose values are read at boot is standard wiring on a Compose stack, not a redesign.

Observability. Finding out that a flow failed means opening the Windmill UI and looking at the run history. Nobody is notified. The hourly schedule limits the damage (the next run catches up on a one-off failure), but a lasting failure, an expired token or an unreachable target, stays invisible until it shows up somewhere else. Windmill’s error handler pointed at an alerting webhook, or a periodic poll of failed runs through its API, are the two obvious answers. Neither requires touching the flows.

What I would tell myself before starting

Decide what the automation may delete, on day one. A sync that only adds is not a sync, it is an import. But one that deletes everything it does not recognise will eventually eat a record you created by hand and forgot. dns_allowed_zones exists because the first version had no such limit.

Make the setup step idempotent. Workspace, variables, flows, schedules, webhooks: all created by a script that can be replayed at will. Anything provisioned by hand becomes the one thing nobody can rebuild.

Keep the flow steps small and separate. Read, compute, apply. Three steps tell you where a run failed in one glance. One step makes you read logs.

Do not publish DNS and DHCP on the host. They are infrastructure for the stack, not services for whatever network the server running the POC happens to sit on. Keeping them unpublished was the easiest security decision in the whole project.

That decision belongs to the Docker Compose POC, but the principle transfers unchanged to production: isolate the network services in a dedicated VLAN and expose only what has to be exposed. In practice that means a DHCP relay rather than a server per VLAN, DoH (DNS over HTTPS) for client resolution, and small read-only DNS servers per zone fed by AXFR. The goal is the same either way: standard behaviour, less network configuration, and certainly not every port open on every VLAN.

The pattern does not stop at DNS and DHCP: each new target is one more read, compute, apply flow. vCenter port groups, switch configurations NetBox already knows how to render from its config templates, firewall address objects (the objects, not the rules, which have to stay under human review). I have implemented none of them, so I will say nothing more about them.

One thing is worth anticipating before getting there: the more consumers there are, the more NetBox data quality becomes critical. With one, a malformed dns_name is one bad record. With six, it propagates everywhere at once.

Share