Why imperative homelab management stops scaling
Imperative infrastructure management works until context switching and state drift make the running estate impossible to reconstruct reliably.

10 min read
On the morning of July 7, 2026, Plex was gone.
It was not stopped or unhealthy. No container by that name appeared anywhere in Unraid’s Docker inventory. The array was up, and 34 other containers were still in the inventory. I suspected an overnight update, but the read-only evidence proved only that Plex was absent. It did not identify the cause or when the container disappeared.
Recovery took one deploy because Plex was declared in a repository. A committed Compose file existed, and a Komodo stack pointed to it. Komodo recreated the container from that declared configuration.
Before I declared Plex in Git, the same problem would have meant reconstructing its configuration. I would have searched for the old docker run invocation, the Unraid template, or whichever half-remembered path mappings and group IDs made hardware transcoding work. Then I would have rebuilt it from incomplete documentation and guesses.
Eight other containers on the same host were stopped. One was configured to start automatically but had not; the others had automatic startup disabled and might have been deliberately dormant. The incident check did not establish their source ownership. Their intended state was another question I could not answer from the inventory alone.
That is my argument for declarative infrastructure. Plex had a known recovery path. For the other containers, I still needed to establish what should exist and why.
What imperative management looked like
In late April 2026, I completed an inventory before starting the migration. I found 11 Docker hosts and roughly 75 containers I could document. About 5% of them existed as declared source in a repository, all on one partially migrated host. Only one host had a Komodo Periphery agent. I reached the rest through SSH, Portainer, or the Unraid web interface, depending on how I had set each service up.
I scored backup coverage by layer. The result was about 4 out of 10 overall. Kubernetes etcd scored zero, Longhorn volumes scored zero, and Docker containers scored two. My notes said, “Complete cluster loss = complete data loss.”
No single reckless decision created that state. It came from reasonable choices made one at a time. A service would seem useful, so I would deploy it. It worked, but the flags that made it work could remain in shell history on one host instead of in maintained source.
I was the real source of truth. Every path into the estate depended on one person’s recollection.
Where it breaks
The threshold is not a specific container count. It is the amount of context needed to explain and recover the services.
The trouble starts when hosts multiply services. One host with 30 containers still has one entry point, one set of conventions, and one place where files tend to live. With 11 hosts and 190 containers, I had 11 contexts to remember. I had to recall which host used Portainer, which one needed Tailscale SSH first, and which one followed the /opt/stacks convention.
The second threshold is quieter. It arrives when what I think is running no longer matches what is running.
I kept working from my mental model while reality drifted away from it. The mismatch surfaced when Plex disappeared and I found eight stopped containers whose intended state I could not explain.
The April inventory showed how deep the gap had become. It did not contain Unraid at all. The storage host running my media stack and roughly 30 containers was missing from the document I had written as the complete picture. I did not notice.
In July, I inventoried the runtime instead of relying on memory. I counted 190 running containers across 11 hosts, plus two hosts I could not read because their agents were unreachable. The April and July totals describe different sets of hosts and use different counting methods, so they are not a growth statistic. They show how much of the estate each method could see.
What the migration cost
Most accounts of GitOps migrations that I found jumped quickly to the benefits. The work before those benefits was the difficult part.
Transcribing the running estate
Nothing improves until the running state has been expressed as configuration. A half-declared stack can be worse than an undeclared one because it creates two places to investigate.
Much of the early work was transcription: reading the running configuration and expressing it as maintained source. The April inventory had roughly 75 documented containers, but documenting a container was not the same as migrating it. The work felt like reverse engineering someone else’s system.
Reworking secret handling
Moving configuration into Git meant that credentials stored in environment variables on individual hosts needed somewhere else to live. I ended up with three layers: 1Password as the vault, SOPS with age for encrypted values that must live in Git, and External Secrets to pull values into Kubernetes. Each layer added another failure mode and another recovery responsibility.
Git history also became part of the security surface. A secret-scanning rollout found a credential in repository history. Fixing the current file was not enough: the exposed credential needed rotation, the source needed correction, and the deployed service needed verification.
Moving configuration into Git did not make secret handling safe by itself. It added places where credentials could accidentally be retained or displayed. Scanning, careful tooling, and a recovery procedure mattered as much as the choice of vault.
Fixing the plumbing
Some work produced no visible improvement. Komodo stacks appeared linked to the repository while the service resolver returned the wrong Compose file. One accepted deploy removed an existing container and failed to recreate it. The committed configuration provided the recovery path.
I also had to pin a CI artifact action three major versions behind because v4 did not work with my Forgejo Actions compatibility mode. A runner registered to my user instead of the organization left jobs unassigned at task_id=0 without a useful error.
Those resolver and compatibility problems took work to resolve. At the end, the services worked as they had before, only now their configuration was declared.
Handling locally built images
Images built only on a host give Renovate no registry version to query for the resulting image. Renovate may still track declared base images and other recognized dependencies, but that does not rebuild or distribute the custom image. A private registry can provide a distribution and versioning path; it also becomes another service to operate.
Depending on the control plane
Forgejo became load-bearing: it held the durable desired state used by the deployment tools. A written recovery target for Forgejo and Komodo was not proof that I could meet it. That required a restore exercise, not another diagram.
Before the migration, losing the Git host would have been inconvenient. Now it sits near the top of the incident tree. I traded many small, familiar failures for fewer unfamiliar failures with larger blast radii.
What it bought me
The Plex recovery was the clearest benefit. The container disappeared, and one deploy brought it back from declared source. I did not have to reconstruct it from memory.
Configuration changes are now diffs I can review. Six months later, I can see what changed and why. Rollback is usually a revert instead of a guess about whether a backup predates the mistake. Commit messages also preserve decisions I would otherwise have to derive again.
Updates became a review queue. In the configuration reviewed in August 2026, Renovate ran nightly against declared versions, refreshed a dependency dashboard, and required approval before opening pull requests. Automerge was disabled. A run verified on August 20 read 244 dependencies from 149 files and exited clean.
A clean run proves that the recorded execution completed. It does not prove that every scheduled run succeeds or that Renovate can see every dependency. Detection depends on a recognized pattern, so an unconventional dependency in an unusual path can remain unnoticed without producing an error. I can verify execution from a receipt, but coverage requires a separate check.
The ownership boundary is also clear. Forgejo holds desired state. Renovate proposes changes. ArgoCD reconciles Kubernetes, and Komodo reconciles Docker. Deployments do not start elsewhere. A running container alone does not prove that it is managed.
The useful distinction is between proposing a change and having authority to apply it. A dependency proposal does not grant permission to alter networking, storage, identity, or other high-impact parts of the estate.
The operating model
There are two reconciliation paths and one source. Renovate proposes dependency changes to Forgejo; it is not a deployment hop. Approved Docker configuration reaches Komodo from Forgejo. Approved Kubernetes configuration reaches ArgoCD from Forgejo.
The tools serve different substrates. ArgoCD reconciles Kubernetes resources through the Kubernetes API. Komodo manages the Docker stacks on participating hosts. Using both lets those workloads keep their appropriate deployment model instead of forcing everything into one orchestrator.
The diagram describes an ownership boundary, not proof of a finished migration. A committed file does not prove that its workload is deployed from that file. Source-to-runtime verification remains separate work.
What I would do differently
Inventory first
The runtime inventory was one of the most useful steps, and I should have done it before planning the migration. Before migrating anything, inspect every host and record what is actually running. You cannot plan the sequence around services you cannot see.
Sort out secrets before services
I moved services and redesigned secret handling at the same time, so I touched several stacks twice. Choose the vault, decide how encrypted values will live in Git, and put scanning in place before moving stacks.
Migrate the control plane later
Bringing up Forgejo and Komodo did not make the control plane fully managed. A better first step would have been to make one unremarkable host fully declarative from end to end. That would have taught me the conventions on something I could afford to break.
Tune dependency automation first
Ungrouped and unscheduled update proposals can create more work than manual updates. I would tune Renovate’s schedule, grouping, and approval rules before allowing it to propose changes across the estate.
Expect a stretch with little visible payoff
Early on, I had absorbed much of the cost without seeing the recovery benefits. Services worked about as well as before, and the main result was a growing repository. That period ended, but it was long enough to make the migration feel like wasted effort.
Why bother at home
“Best practice” is not a sufficient reason at this scale. Nobody pays me for homelab uptime.
I would rather spend my evenings on work other than reconstructing old commands. Imperative management postpones that reconstruction until something breaks.
The less practical reason is probably the most honest one. I wanted to answer, “What is running in my house, and why?” The dated inventories showed how incomplete my answer had been. Declared configuration gave me a better place to start, and that has been worth the work.
The figures in this article come from dated inventories of my own estate and describe points in time. They do not reflect current addressing, credentials, or configuration.
