Kubernetes platforms that stay supported, survive a node failure, and upgrade without an outage.
Anyone can build a cluster. The problem starts afterwards: the version drifts away from the supported window and the upgrade gets postponed, because nobody in the team will guarantee that what runs on top survives it. After two years of postponing, the platform that was meant to speed up delivery is the most fragile thing in the environment.
Two questions run through everything here.
Findings ranked by severity — configuration, security baseline, resource governance, persistent storage. How far the platform is from a supported version, whether monitoring actually sees nodes, pods and volumes, and what a single node going down would take with it.
Whether a proven change path exists — a parallel environment and a switchover, rather than surgery on a running platform. How many applications survive a node being drained. And whether certificates renew themselves: with mTLS in place, an expired certificate is a full outage.
Distance from a supported version is the one metric you can check yourself in ten minutes. It doesn't grow evenly, either — fall behind several versions and you are no longer facing one upgrade, but a chain of them.
A platform goes through three states: build it, keep it healthy, change it.
Build it
Nodes become a platform: the distribution and the infrastructure a cluster is unusable without, the networking and persistent storage under it, and the first application packaged onto it and running.
Keep it healthy
The platform stays operable: findings closed, tenants and resources under governance, service identity and certificates handled automatically, patching kept current — and integration problems between platform, storage and network resolved before an application feels them.
Change it
The platform moves to the next version without an outage. We build a parallel environment and switch the load over instead of operating on the running one; the original stays complete until the new one proves it works. It costs capacity for two clusters while the transition runs.
What we deliver
Four ways to engage
The same lifecycle as everywhere at G2F — assessment, build, operations, enablement — focused on one question: is the platform in a state where it can still be changed?
01
AssessKubernetes Health Check & Platform Assessment
An independent read of what state the platform is in — and whether it can still be changed.
- Inventory and findings by severity — cluster configuration and layout, security baseline, resource governance, persistent storage, and how Helm charts are managed.
- Control plane and recoverability — cluster state backed up, and a tested restore. The usual answer is that it needs no backup because the platform can be rebuilt from its declaration — true only when the declaration is complete, including whatever was pushed by hand.
- Distribution fit — vanilla, lightweight with multi-cluster management, security-hardened for regulated environments, or enterprise with commercial support. Compared against licence model, vendor lock-in and the operational footprint against the team that will hold it.
- Upgrade path and its blockers — the real blocker is rarely Kubernetes itself: third-party images that stopped being freely available or changed licence terms, and components with nowhere to go.
- A prioritised roadmap — findings ranked by importance and effort, with a short summary in the language of a board.
02
BuildFrom installation to a platform you can run
A platform your own team can operate and move forward — built from nodes, or brought back into a changeable state.
- Installation and what sits under it — the distribution plus what a cluster is unusable without: load balancing (bare metal included), DNS, image registry, storage classes with CSI, CNI and network policies.
- Applications onto the platform — analysis, Helm charts, database and cache operators with backup and restore, and the first deployment. New versions afterwards belong to the change path.
- Upgrade without an outage — a parallel environment and a switchover. Images with nowhere to go are handled before the upgrade, not during it; the original stays complete until the new one proves it runs.
- Service mesh, module by module — onboarded in small groups, with mTLS and an automated certificate lifecycle. The certificate authority becomes a critical path, and onboarding stays reversible per application.
- Multi-tenancy — namespaces, quotas and policies for teams inside one organisation. Mutually untrusted tenants need node or cluster separation, which is a different delivery.
- Before handover — node-drain behaviour is tested and a way back from the upgrade exists. The platform moves the load; an application running in a single replica falls over regardless.
03
OperatePlatform operations & on-demand
A platform you build and forget goes backwards — versions move quickly, and integration problems only surface in production.
- Proactive platform monitoring, L2/L3 support, and on-demand consulting blocks with direct access to the consultant.
- Version maintenance and patching, so the distance from a supported version stops growing, plus capacity and configuration changes as the platform grows.
- Integration troubleshooting after go-live — MTU mismatches between pods, containers and the gateway, traffic routed the wrong way, volumes hanging because CSI traffic crosses a firewall, database cluster recovery.
- We take over someone else's platform after a health check, not blind — without a known starting point you inherit the findings along with the platform, and the argument over what counts as an incident arrives in month one.
04
TrainKubernetes workshops for your team
Your team should be able to hold the platform after we leave.
- Container images — permissions and handling secrets, layering and multi-stage builds.
- Kubernetes networking — CNI, services, ingress and network policies.
- Modular and built around your environment. Or hands-on inside a delivery, where your team takes the platform over as it is built.
How we come in
We arrive without a preferred platform. Which distribution fits is an output of the assessment, not a premise we bring with us — measured against your environment: licence model, vendor lock-in, how much of the platform sits behind a vendor abstraction, and the operational footprint against the size of the team that will run it. Where commercially backed vendor support is what your organisation requires, that is what we deploy. Where an open-source-first policy or a public tender rules it out, we deploy the equivalent open-source layer. We are not a reseller of any distribution — the commercial relationship with the vendor stays yours.
Delivered
Where we've done this
Platforms we've assessed, built, upgraded and run — for regulated fintechs, banks, system integrators and critical infrastructure operators.
Three applications containerised into Helm charts, with database and cache operators including backup and restore — a platform running in production at a regulated fintech.
Read the case study →
A long-term platform operations partnership — proactive monitoring, ongoing upgrades, and distributed storage behind persistent volumes.
Read the case study →
🇨🇿
A service mesh estate analysed and split into three streams for phased adoption — with mTLS, an automated certificate lifecycle and a governance model for a bank.
Read the case study →
An enterprise platform installed on bare metal for a system integrator, with the load balancing, DNS, registry and storage layer underneath it — plus a series of Kubernetes workshops.
Platform selection advisory for a critical infrastructure operator — a distribution comparison, a concept design with automation, and a working demo.
How far are you from a supported version — and what happens if you drain a node?
A Kubernetes Health Check answers both in a fixed scope: findings ranked by severity, the gap to a supported version, and an upgrade path with whatever blocks it. The result is ordered so the fixes can be packaged straight into work. Most first conversations take 30 minutes. No pitch, no deck.