
One year of Talos in production
A field report after twelve months of Talos Linux in production, in the homelab and on throwaway single node clusters: Terraform-only provisioning, upgrades with no outage, what a read-only rootfs actually buys, and the two places it still hurts.
A year ago the cluster ran on general purpose Linux nodes: a distribution, a configuration management tool, an SSH bastion, an on-call runbook full of commands to type on a node at three in the morning. Today it runs on Talos, and the honest summary of the year is that the operating system stopped being something I think about.
That sentence is easy to write and hard to earn, so what follows is the detail: what the cluster actually looks like, what a year of upgrades and failure tests produced, and the two things I would still like fixed.
The production cluster is the demanding case and it gets most of the space, but it is not the whole year. The same Talos runs in my homelab, and on the single node clusters I create to evaluate a piece of software and delete the same evening. That those three settings are one operating system, one tool and one configuration format turned out to be a finding in itself, so it gets a section of its own further down.
On this page
The cluster
Baremetal, three sites, one Kubernetes cluster.
site A site B
┌────────────────────────────────┐ ┌────────────────────────────────┐
│ cp-01 control plane │ │ cp-02 control plane │
│ │ │ │
│ worker-01 .. worker-05 │ │ worker-06 .. worker-10 │
│ (Ceph OSDs + workloads) │ │ (Ceph OSDs + workloads) │
└───────────────┬────────────────┘ └───────────────┬────────────────┘
│ │
│ L3 between sites │
└───────────────┬───────────────────────┘
│
┌────────────┴─────────────┐
│ site C │
│ cp-03 control plane │
│ no workers, no data │
│ third man for quorum │
└──────────────────────────┘Site C holds nothing but a control plane node. It exists so that losing site A or site B leaves two etcd members out of three, which is a majority, which means the API server keeps answering. It is the cheapest piece of the whole design and the one that decides whether a site outage is an incident or a non event.
Each worker is a two socket server with thirteen disks, and the split matters more than the count:
worker-0X
┌──────────────────────────────────────────────────────────────────────┐
│ 2 x SSD, hardware RAID 1 ─▶ Talos system disk, read-only rootfs │
│ 1 x NVMe ─▶ Talos user volume, xfs, local-path PV │
│ 1 x NVMe ─▶ Ceph OSD, device class `nvme` │
│ 1 x NVMe ─▶ Ceph metadata (RocksDB/WAL) │
│ 4 x SAS 10k ─▶ Ceph OSD, device class `sas-fast` │
│ 4 x SAS 7.2k ─▶ Ceph OSD, device class `sas` │
└──────────────────────────────────────────────────────────────────────┘
13 disks, every one of them selected by its /dev/disk/by-path entryThree things about that layout, and only the first is really about disks.
Every disk is declared by its bus path, never by /dev/sdX or
/dev/nvme0n1. Kernel enumeration order is not a contract: it can change
after a reboot, a firmware update or a backplane swap. On a node whose whole
identity is a configuration file, that is not a cosmetic problem, it is the
difference between “install on the RAID 1 volume” and “install on an OSD”. The
system disk is matched on busPath in the Talos configuration, and the Ceph
disks are handed to Rook through a devicePathFilter on the same
/dev/disk/by-path entries. Take the ten minutes to write the bus paths
down. It is the single cheapest guarantee in the whole build.
The system disk is a hardware RAID 1 volume, on purpose. Talos has no package manager and no shell, so a system disk failure is not something you fix by logging in. The controller hides it, the node keeps running, and the disk is swapped in business hours rather than during an incident. Nothing of value lives there anyway: the machine configuration is in Git.
The local disk is a Talos user volume, formatted xfs by Talos itself, and consumed by local-path-provisioner. That is a deliberately temporary answer to “some workloads want a fast local PV and no replication”, and it will be revisited. It is worth noting because it is the one piece of storage Talos provisions directly: everything Ceph touches is declared nowhere in the Talos configuration, left raw, and claimed by the storage operator. Declaring an OSD disk in both places is how you get a partition table that two systems disagree about.
The Ceph side has one design decision worth flagging and then leaving alone: metadata (RocksDB and WAL) sits on a dedicated NVMe to keep the SAS OSDs fast, which makes that single device a shared failure domain for every OSD behind it. It buys real performance and it is on the list to redo. It is also a Ceph question rather than a Talos one, so it stops here.
Provisioning: Terraform, and that is the whole story
The cluster is a Terraform state and a directory of Jinja2 templates. There is no Omni, no PXE orchestrator with its own database, no configuration management agent running on the nodes.
resource "talos_machine_secrets" "this" {}
data "talos_machine_configuration" "this" {
for_each = var.nodes
cluster_name = var.cluster_name
cluster_endpoint = "https://${var.api_vip}:6443"
machine_type = each.value.role # controlplane | worker
machine_secrets = talos_machine_secrets.this.machine_secrets
talos_version = var.talos_version
# one rendered patch per concern, applied in a fixed order
config_patches = local.patches[each.key]
}
resource "talos_machine_configuration_apply" "this" {
for_each = var.nodes
client_configuration = talos_machine_secrets.this.client_configuration
machine_configuration_input = data.talos_machine_configuration.this[each.key].machine_configuration
node = each.value.ip
}local.patches is where the rendering step lands: for every node, each template
rendered against that node’s facts, which are its role, its zone, its addresses
and its bus paths. The templates are split by concern, never by machine:
patches/
├── network.yaml.j2 # addresses, VIP, bonding
├── dns.yaml.j2
├── ntp.yaml.j2
├── proxy.yaml.j2
├── registry.yaml.j2 # mirrors, per site
├── disks.yaml.j2 # system disk, user volumes
├── kubelet.yaml.j2
└── extensions.yaml.j2That split is the part I would keep in any future cluster. One file per
concern means “change the NTP servers” is a one file diff, reviewable by
someone who knows nothing about the rest of the configuration. And where Talos
exposes a real document kind for a concern, the patch is that document rather
than a v1alpha1 fragment: UserVolumeConfig for the local volume,
ExtensionServiceConfig for extension settings, NetworkRuleConfig for the
ingress firewall. Those are validated on their own, so a mistake surfaces at
apply time instead of being merged silently into a nested map.
The differences between roles and between sites live inside the templates, as conditionals, not as duplicated files. The disk patch is the clearest example, since it is also where the layout above becomes machine readable:
# patches/disks.yaml.j2
machine:
install:
# matched on the bus path, never on /dev/sda, which is only true
# until the next reboot
diskSelector:
busPath: '{{ host.system_disk_bus_path }}'
{% if host.role == 'worker' %}
kubelet:
extraMounts:
- destination: /var/mnt/local-path
type: bind
source: /var/mnt/local-path
options: ['bind', 'rshared', 'rw']
---
apiVersion: v1alpha1
kind: UserVolumeConfig
name: local-path
provisioning:
diskSelector:
match: disk.bus_path == '{{ host.local_disk_bus_path }}'
filesystem:
type: xfs
{% endif %}Talos formats and mounts that volume itself, under /var/mnt/local-path. The
extraMounts entry is what lets local-path-provisioner hand it to a pod. A
control plane node renders none of that block: it has two disks and nothing to
provision.
Everything else, the ten Ceph disks included, is deliberately absent.
The site axis works the same way, and site C is the interesting case precisely because it is not a data site:
# patches/registry.yaml.j2
machine:
registries:
mirrors:
docker.io:
endpoints:
- https://mirror.{{ host.zone }}.example.internal
{% if host.zone == 'dc3' %}
# no local mirror there: quorum site, no workloads to serve
- https://mirror.dc1.example.internal
{% endif %}Two axes, role and zone, and keeping it to two is the whole discipline. Six
full configurations, one per role and per site, would have drifted within a
quarter. Two axes of conditionals inside one template per concern do not,
because the shared part is physically shared rather than copied and kept in
sync by good intentions.
The cost is real: you can no longer read a node’s configuration off the disk,
only the template that produces it. The answer is
talosctl apply-config --dry-run, which prints the diff against what the node
is currently running. Render, diff, then apply. That habit is what keeps
templated configuration honest.
What this buys is not “infrastructure as code” as a slogan. It is a specific, checkable property: the machine configuration is the whole node. There is no drift to reconcile because there is no second place a change can be made. No one has ever fixed a node by hand here, because there is no way to.
Omni is a good product and I understand why it exists: it solves discovery, enrolment and access at once, for fleets that are not declared in a Terraform state. Ours is. Adding a SaaS control plane, or self-hosting one, to manage thirteen machines already described in Git would have been a second source of truth bought at the price of the first.
Upgrades: twelve months, zero outage
Node upgrades:
talosctl -n worker-07 upgrade \
--image factory.talos.dev/installer/<schematic>:v1.x.yKubernetes upgrades:
talosctl -n cp-01 upgrade-k8s --to 1.3x.yThe second command is the one that changed how the team feels about Kubernetes minor versions. It walks the control plane components and the kubelets in order, one node at a time, and it is a single command rather than a runbook.
Across the year, node upgrades and Kubernetes upgrades produced no outage and no degraded window that a user noticed. The operational feel is the one you get from EKS or AKS: you pick a version, you trigger the upgrade, you watch it roll. The difference is that the nodes are ours, in our racks, and the control plane is not a black box we file a ticket against.
Two things make that true, and neither is luck:
- The upgrade is an A/B image swap, not a package transaction. A node either comes back on the new version or rolls back to the previous one. There is no half upgraded state to debug, which is exactly the state that turns a Linux distribution upgrade into an incident.
--stageexists for the cases where a workload cannot be drained cleanly. The upgrade is written to disk and applied on the next reboot, which lets the reboot happen in a window you chose.
The read-only rootfs, and what it actually covers
A local privilege escalation in the kernel or in a userspace component made the rounds recently. The cluster was not exposed, and it is worth being precise about why, because “immutable OS” gets used as a magic word.
Talos removes the exploitation path, not the bug:
- there is no shell and no SSH daemon, so there is no interactive session on the node to escalate from;
- there is no package manager and no writable
/usr, so a dropped binary has nowhere to persist; - the root filesystem is read-only and verified, so tampering with a system binary is not a matter of permissions but of the filesystem refusing writes;
- the only interface is the machine API, mutual TLS, with a role on the client certificate.
What it does not do is patch the kernel. A remote code execution reachable from a network service, or a container escape, remains something you fix by upgrading to a Talos release that carries the fixed kernel. The difference is that this upgrade is the command from the previous section, applied on a Tuesday, instead of a distribution upgrade coordinated across thirteen machines.
The honest framing: Talos does not make you immune, it makes the blast radius of a whole class of local vulnerabilities empty, and it makes the fix routine.
Failure tests, all of them passed
Everything below was tested deliberately, on the production cluster, during maintenance windows:
| Test | Result |
|---|---|
| Kill one worker, hard | Pods rescheduled, Ceph recovered, node rejoined on boot |
| Reboot a control plane node | etcd member rejoined, no API interruption |
| Full site outage, site A | Quorum held on cp-02 and cp-03, workloads moved to site B |
| Inter-site link cut | Minority side fenced, majority side kept serving |
| Rolling reboot of all workers | No workload downtime, no manual step |
Nothing on that list required a human to type anything on a node. In every case the recovery was: power comes back, the node boots, it fetches its configuration, it rejoins. A Talos node has no state worth preserving between reboots, so “did it come back correctly” stops being a question.
One note the tests surfaced, and it is not a Talos problem: etcd quorum is not the only quorum. Storage spread over two data sites has its own, and it has to be designed alongside the control plane layout rather than after it.
The one thing I want fixed: access to talosctl
This is my only real complaint after a year, and what sharpens it is that the same cluster already solves the identical problem one layer up.
The Kubernetes API server authenticates against Entra ID over OIDC. Access
is a group membership: someone joins the on-call rotation, they land in the
right group, their kubectl works, and it stops working the day they leave.
Nothing is distributed by hand, nothing is revoked one certificate at a time.
The Talos machine API is the layer underneath, and it works differently. Access
is a client certificate carrying a role (os:admin, os:operator,
os:reader). The permission model is sound, and it is a genuine improvement on
“everyone with an SSH key is root”. What it lacks is that same bridge to the
identity provider, so one cluster ends up with two access stories: Kubernetes on
SSO, and the machines under it on certificates that get copied around.
On classic Linux nodes this is a solved problem too: an LDAP group, a sudo
rule, an SSH certificate authority, and the person who joined the rotation last
week gets access by being added to a group. There is no equivalent binding in
open-source Talos. Omni provides SSO, but that means adopting Omni for an access
control feature, which brings back the second source of truth discussed above.
What we do instead works, but is more machinery than it should be: a short-lived operator certificate, issued on demand rather than distributed.
talosctl config new /tmp/oncall.talosconfig \
--roles os:reader \
--crt-ttl 8hIt is not group membership, it is not centrally audited, and someone has to run it. The gap is not the permission model, it is the enrolment.
Someone has already built the missing piece:
talosctl-oidc, by Quentin Joly, which
is written up here. The shape
is exactly what the problem calls for. A server holds the Talos CA key and
authenticates the user against the identity provider, OIDC claims are mapped to
Talos roles (a platform-admins group to os:admin, a developers group to
os:reader), and the client exchanges an ID token for an ephemeral certificate,
five minutes by default, written straight into the talosconfig. It automates the
command above and then does the part the command cannot: deriving the role from
a group rather than from whoever typed it.
Two things to weigh before putting it under an on-call rotation. It is young, and it holds the CA private key, which makes it a component to protect like a certificate authority rather than like a helper binary. Neither is a reason to ignore it. It is the right answer to the right problem, and I would rather see this shape land upstream than keep maintaining the workaround.
The other rough edges: hardware, and not really Talos
Two areas took real time this year, and in both cases I think Talos is the messenger rather than the cause.
Wiping a disk. Reinitialising an OSD or a whole storage cluster is awkward,
because the usual reflex is a shell on the node and sgdisk --zap-all, which
does not exist here. What works is a privileged Job with the device mapped in,
or talosctl wipe disk on recent Talos versions. Both are fine once written
down. The friction is that nearly every procedure published online assumes a
node you can log into, so it needs translating before it can be used.
GPUs. Drivers on Talos are system extensions baked into the boot image through the Image Factory, which means the driver version is pinned to the Talos version by a schematic. Upgrade Talos and you rebuild the schematic; the kernel module, the container toolkit and the device plugin all have to agree. When the vendor ships an extension for your card, this is clean and reproducible, better than a driver installed by a DaemonSet at runtime. When it lags, or when the card is exotic, you wait. That is a vendor and ecosystem problem that happens to surface at the Talos boundary.
Neither of these made me reconsider the choice. Both are worth knowing before you sign up for a storage-heavy or GPU-heavy cluster on an immutable OS.
The same thing in the homelab and on throwaway clusters
A good part of the year’s Talos usage was neither production nor thirteen machines. It was the homelab, and single node clusters brought up to try software out and torn down once the question was answered.
A single node cluster is the same configuration with one line changed:
cluster:
allowSchedulingOnControlPlanes: trueSame templates, same talosctl, same upgrade command, same patches split by
concern, with a role axis that happens to have one value. Nothing about the
workflow changes at one node, and that is the actual point: what you learn on
a throwaway cluster transfers, because it is not a different system. A
distribution installed by hand on a test VM always drifts from what production
runs, and every conclusion drawn on it carries an asterisk.
The other half of the value is how cleanly it goes away. There is no state to
clean up and no host left half configured: talosctl reset returns the machine
to nothing, or the VM is deleted and nothing on it mattered. Evaluating three
storage operators in a week stops being an exercise in undoing the previous two.
What one node does not tell you is worth stating plainly. Quorum behaviour, failure domains, and how a rolling upgrade interacts with pod disruption budgets are all properties of having several machines. The single node PoC validates the software. It does not validate the cluster.
My conclusion
Talos was the right call, and one reason carries most of the weight.
A node no longer has a history. Nothing was installed by hand during a debugging session and left behind. No one set a sysctl in 2024 that nobody can explain today. The first node built and the last one added run the same thing, because both are the output of the same configuration file.
Everything else in this note follows from that. Upgrades are boring. Failure tests pass. An incident never starts with “can you get me on the node”.
The effort is the same on baremetal and in a VM, and it stays small. There is no technical debt to carry, because there is nowhere for it to accumulate.
The two complaints above still stand. talosctl access needs a real identity
provider, and some hardware still asks for patience. Neither is a reason to go
back.

