Skip to main content
Version: Developer

(RA-K1) Kubernetes Control Plane with External Docker Agent Architecture Explained

The RA-K1 architecture is usually described as "RA2, but the control plane runs in Kubernetes." That description is accurate about the topology. It is misleading about what actually changes.

RA2 separates decision-making from execution by putting them on different servers. RA-K1 keeps that same separation and adds a second, less obvious one: the control plane and the execution plane stop sharing a lifecycle model. In RA2, both sides of the boundary are virtual machines, installed the same way, patched the same way, recovered the same way. In RA-K1, one side is a set of declarative Kubernetes workloads that an orchestrator continuously reconciles toward a desired state, and the other side is a virtual machine that nothing reconciles at all.

This is not a packaging change. It is the introduction of a second orchestrator into the system, and almost everything interesting about RA-K1 follows from the fact that the two orchestrators cannot see each other.

note

This page documents Kasm Workspaces 1.19.0 and the Kasm Helm chart 1.1190.x (oci://registry-1.docker.io/kasmweb/kasm-helm). Later chart versions may rename specific values or templates.


Reference Architecture K1​

Kasm Workspaces RA-K1 Kubernetes control plane and external Docker Agent architecture diagram

Click the diagram to open it full size.


The Primary Load Balancer​

Every Kasm deployment fronts its services with an NGINX component, kasm_proxy. In RA-K1 that component is still present, running as a Deployment in the cluster, but it is no longer the first thing a user's packet touches. In front of it sits the Kubernetes ingress path, and in front of that sits the Primary Load Balancer.

The chart exposes the deployment in one of three ways:

Exposure methodChart settingNotes
Service of type LoadBalancerproxyService.type: LoadBalancer (default)The cloud provider provisions an external L4 load balancer directly in front of the kasm_proxy pods. TLS is terminated by kasm_proxy inside the cluster.
Ingress resourceingress.enabled: trueAn ingress controller handles L7 routing and, typically, TLS termination. Requires an ingressClassName and a controller already running in the cluster.
OpenShift Routeroute.enabled: trueThe OpenShift-native equivalent of the ingress path.

Whichever path is chosen, the architectural role is identical to the Primary Load Balancer in RA1 through RA4: a single public endpoint through which authentication, session initiation, and workspace streaming all enter the environment. The deployment presents itself to users as one service with one name, regardless of how many pods and how many virtual machines are actually behind it.

Organizations with mature security tooling should place an enterprise load balancer, WAF, or Zero Trust access proxy in front of the ingress path rather than treating the cluster edge as the security boundary. Web Application Firewall inspection, DDoS mitigation, and TLS offload all belong at this tier, where policy can be applied uniformly to every flow before it reaches any Kasm component.

For production deployments, a network load balancer with active health checking is the recommended approach. The App role exposes /api/__healthcheck, which verifies both application liveness and database reachability. An api pod that cannot reach PostgreSQL reports unhealthy and is pulled from rotation. Note carefully what that means in RA-K1: when the database is down, every replica reports unhealthy at the same time, and the load balancer removes all of them. Health checking protects against a single sick replica. It does not protect against a sick dependency.

warning

A single-replica ingress controller silently reintroduces the RA1 failure mode. If the entire control plane is running at deploymentSize: large with three replicas of every component behind one ingress controller pod, the deployment has one point of failure, not none. Replicate the ingress controller, or terminate at an external load balancer that is itself redundant.


The Defining Characteristic: Two Orchestrators, One System​

In RA1, all roles collapse onto one machine and there is no boundary. In RA2, roles separate across machines and the boundary is a network hop. In RA-K1, the boundary is a network hop and a change in how workloads come into existence.

Inside the cluster, Kubernetes owns the control plane. It decides which node runs kasm_api, restarts it when it dies, reschedules it when a node is drained, and continuously drives the running state toward what the Helm values declare. None of this is Kasm-specific. It is the standard Kubernetes contract.

Outside the cluster, Kasm owns the execution plane. kasm_manager decides which Agent will host a session, the Agent instructs its local Docker Engine to start a container, and the container runs until the session ends. None of this is Kubernetes-specific. Kubernetes has no visibility into it whatsoever.

The result is a system with two independent schedulers operating on two different populations of containers, separated by a boundary that neither can cross:

  • Kubernetes schedules pods. It knows nothing about Kasm sessions, cannot see the Agent VM, and will not react to anything that happens there.
  • Kasm schedules sessions. It knows nothing about nodes, pods, or PodDisruptionBudgets, and treats the entire cluster as a single logical control plane.

Teams that come to Kasm from a Kubernetes background arrive with a strong prior that the platform is self-healing, because in their experience everything on a cluster is. In RA-K1 that prior is correct for the control plane but wrong for session compute.


What the Helm Chart Actually Deploys​

Kasm is composed of four service roles. The Helm chart deploys three of them and deliberately excludes the fourth.

RA1 through RA4 name the control plane a single "Web App role" and never distinguish its containers. The Helm chart cannot do that: Kubernetes schedules, scales, and heals Deployments, not roles, so what those pages call the Web App role becomes two separate objects here — an App Deployment (kasm_api, kasm_manager) and its own Dedicated Proxy Deployment (kasm_proxy).

RoleContainsDeployed by the chartKubernetes workload type
Appkasm_api, kasm_managerYesDeployments
Dedicated Proxykasm_proxy (NGINX)YesDeployment
Connection Proxykasm_guac, kasm_rdp_gateway, kasm_rdp_https_gatewayYesStatefulSets (guac, rdp_https_gateway) and a Deployment (rdp_gateway)
DatabasePostgreSQLYes, optionallyStatefulSet, always one replica
AgentDocker Engine, kasm_agent, kasm_proxyNoNot a Kubernetes workload at all

The deploymentSize value (small, medium, large) sets CPU and memory requests and limits across all components and scales replica counts to 1, 2, and 3 respectively. Individual components can be overridden through components.<name>.replicas and components.<name>.resources where the preset does not fit.

The database StatefulSet and rdp-gw are the exceptions. They run a single replica at every deployment size.

** Choosing large produces a horizontally scaled, genuinely resilient control plane in front of a single-instance database and rdp-gw. This is covered in detail below, because it is the property of RA-K1 most likely to be misread as "we deployed it HA."


Why the Agent Lives Outside the Cluster​

The most common question about RA-K1 is why the Agent is not simply another Deployment. The chart could package it. It deliberately does not, and Kasm's guidance is explicit that Agents are always deployed on standard Docker hosts outside Kubernetes.

The reasons are architectural rather than incidental:

The Agent manages a container runtime. It is not managed by one. kasm_agent talks directly to the Docker Engine on its host to create, inspect, and destroy session containers. Running that inside a pod means running a container runtime inside a container runtime, with the privilege escalation and lifecycle ambiguity that implies. The Agent is not a workload that happens to use Docker. The Agent is a Docker host with a control interface attached.

Session containers need capabilities that pod security models exist to prevent. Workspace images require specific /dev/shm sizing, seccomp profiles tuned for desktop workloads, device passthrough for GPU-accelerated workspaces, and in some cases privileged operations. These are exactly the properties a hardened Kubernetes cluster is configured to refuse.

Kubernetes rescheduling semantics are actively harmful for interactive sessions. The behavior that makes Kubernetes excellent for stateless services (evict a pod, reschedule it elsewhere, let the Service route around it) is destructive when the pod is somebody's live desktop with unsaved work in it. A node drain during routine maintenance would terminate every session on that node. Kasm's session lifecycle assumes a session ends when the user ends it, not when the scheduler decides to rebalance.

The capacity model is VM-shaped. Kasm sizes Agents by concurrent sessions per host, with fixed CPU and memory reservations per workspace. Agent AutoScaling provisions and deprovisions virtual machines through a cloud provider. This maps cleanly onto a VM pool and awkwardly onto a pod population competing with the control plane for the same node resources.

The result is a clean division of concerns: put the stateless, reconcilable, request-response workloads where Kubernetes is strongest, and keep the stateful, long-lived, privilege-hungry, latency-sensitive workloads on infrastructure that models them naturally.

note

The Agent is required for any zone that runs container sessions, including the primary zone. A Helm install with no Agent produces a fully functional administrative UI in which no containerized workspace can ever start. RDP, VNC, and SSH sessions work immediately, because those flow through the Connection Proxy that the chart does deploy.


Four Traffic Flows Define the Architecture​

Reading RA-K1 as a box diagram is misleading. Reading it as a set of flows that diverge after session creation is accurate.

Kasm Workspaces RA-K1 traffic flows diagram: control plane, execution, session plane, and remote system flows

The control plane flow carries authentication, policy evaluation, and orchestration decisions. Browser to Primary Load Balancer, to ingress, to a kasm_proxy pod, to a kasm_api pod, to PostgreSQL. This flow is entirely contained within the cluster once it passes the edge. It is the only flow that touches the database, and it is the flow that stops working first when anything is wrong.

The execution flow carries session provisioning instructions. The kasm_manager pod, through cluster egress, to kasm_agent on the VM, to the Docker Engine, to a running container. This flow crosses the cluster boundary outbound, on port 443, to the Agent's registered hostname. It is the flow most often broken by network policy, because it is the one that requires pods to reach a destination that is not in the cluster.

The session plane flow carries the pixel stream. Browser to Primary Load Balancer, to a kasm_proxy pod, to kasm_proxy on the Agent VM, to the session container over KasmVNC. This flow leaves the cluster and stays out. It is database-independent once established, which produces the counterintuitive behavior that a control plane outage does not immediately drop a streaming session. The session simply cannot be re-authenticated, and it dies at token renewal.

The remote system flow carries legacy protocols. Browser to Primary Load Balancer, to a kasm_guac pod, to the target system over RDP (3389), VNC (5900), or SSH (22). No Agent is involved. This flow is what makes RA-K1 useful on day one, even before any Agent exists.

The separation is the point. A user streaming a containerized desktop and a user streaming an RDP session to a Windows server are traversing almost entirely different infrastructure, sharing only the authentication path and the identity that produced the token.


Two Session Types, One Control Plane​

RA-K1 supports the same two session types as every other Kasm architecture.

Containerized workspaces are provisioned on the external Docker Agent and delivered as ephemeral containers. Their capacity is a function of Agent VM count and size, a dimension that has nothing to do with the Kubernetes cluster. Adding cluster nodes does not add workspace capacity. Adding Agent VMs does.

Remote system access sessions connect to existing Windows, Linux, and macOS systems over RDP, VNC, or SSH, mediated by the Connection Proxy running as pods in the cluster. Their capacity is a function of cluster resources, because kasm_guac is doing the protocol translation and it is running on cluster nodes.

Both share the same authentication flow, the same control plane, and the same public hostname. They scale along entirely independent axes. An organization that grows its containerized workspace usage will exhaust Agent VMs while the cluster sits idle. An organization that grows its RDP usage will exhaust kasm_guac pods while the Agent VMs sit idle. Capacity planning in RA-K1 requires modeling both curves separately.


The Cluster Boundary Is a Trust Boundary​

RA-K1's structure enforces Zero Trust properties architecturally rather than by configuration.

  • The session container is never directly reachable by the user. The only path is Primary Load Balancer, to a kasm_proxy pod, to the Agent's kasm_proxy, to the container, and every hop validates the session token against the control plane.
  • The Agent VM has no public interface. It accepts connections only from the cluster's egress addresses on 443, and it initiates connections only to the deployment's public hostname.
  • The database has no public interface. Only kasm_api and kasm_manager reach it, on 5432. If the bundled StatefulSet is used, it is reachable only through a ClusterIP Service that never leaves the cluster network.
  • RDP, VNC, and SSH are terminated at kasm_guac and converted to WebSocket before anything reaches the browser. Those protocols are never exposed to the internet.
  • The Primary Load Balancer is the single controlled ingress point where inspection and WAF policy apply uniformly.

Kubernetes adds an enforcement mechanism the VM-based patterns do not have: NetworkPolicy. The pod network can be constrained so that only kasm_api and kasm_manager may egress to 5432, only kasm_manager may egress to the Agent's address, and nothing else may leave the namespace at all. This is worth doing. It is also worth remembering that the strongest NetworkPolicy in the world stops at the cluster edge. The Agent VM's own firewall is a separate artifact that must be maintained separately, and it is the most commonly neglected control in this pattern.


Identity and DNS: Three Names That Must Resolve​

RA-K1 has one public hostname and two additional resolution requirements.

NameSet byMust be resolvable and reachable byCommon failure
publicAddrHelm valuesEnd users, and the Agent VMSet to an internal-only name. Users cannot reach it, or the Agent cannot check in.
MANAGER_HOSTNAMEAgent install flagThe Agent VM, outbound on 443Set to a pod IP or a .svc.cluster.local name, unreachable from outside the cluster and unstable even inside it
AGENT_HOSTNAMEAgent install flag (--public-hostname)The kasm_manager pods, outbound on 443Set to a private IP that the pod network has no route to, or to a name only the VM's own resolver knows

The Agent's manager hostname must be the deployment's stable public address, not anything cluster-internal. Pod IPs change on every restart. ClusterIP Services do not resolve outside the cluster. The Agent is outside the cluster. The Agent must therefore reach the control plane the same way a browser does, through the Primary Load Balancer, at publicAddr, on 443.

The reverse direction has the mirror requirement. kasm_manager runs in a pod and must open a connection to AGENT_HOSTNAME on 443. Whatever value the Agent reported at registration is the value the pods will use, so it must be a name or address routable from the pod network. In practice this means the Agent VM sits in a subnet the cluster's node network can reach, and AGENT_HOSTNAME is its private DNS name or private IP.

TLS follows the same shape as every other pattern: the certificate must match publicAddr, WebSocket origin must match, and the Proxy Hostname configured in the zone must resolve identically. The chart accepts a pre-created Kubernetes secret through certificate.secretName, or drives cert-manager through certificate.certManager.

warning

If Agent AutoScaling is enabled later, the Upstream Auth Address zone setting must point at the load balancer address, not at any individual pod or node. Newly provisioned Agent VMs that cannot reach the App role's API fail to register silently: the cloud provider creates the VM, Kasm never sees a check-in, and capacity does not increase. This is the single most common AutoScaling failure mode.


The Database: The One Thing Kubernetes Does Not Make Resilient​

The database is the only stateful component in the deployment and the single most consequential availability dependency. Its role in RA-K1 is exactly what it is in RA1 and RA4: the source of truth for users, sessions, configuration, and authentication state. Nothing about running on Kubernetes simplifies it.

The bundled kasm_db StatefulSet runs one replica at every deploymentSize. Kubernetes restarts the pod if it dies and reattaches its PersistentVolumeClaim. That is a restart, not a failover. There is no standby to promote and no replication, and recovery time is bounded by how long PVC reattachment and PostgreSQL crash recovery take. During that window the entire platform is down, because every kasm_api replica is failing its health check against the same unreachable dependency.

Redundancy approachMechanismSuitable for
Managed cloud PostgreSQL (RDS Multi-AZ, Cloud SQL HA, Azure Database)Synchronous standby, automatic failover, managed backupsThe recommended choice for any production RA-K1 deployment
Aurora PostgreSQLClustered storage, fast failoverHigh-throughput cloud deployments
Self-managed PostgreSQL with Patroni, outside the clusterStreaming replication with automated failoverOn-premises and private cloud
A PostgreSQL operator inside the clusterOperator-managed replication and failoverTeams with existing operational maturity running databases on Kubernetes
Bundled kasm_db StatefulSetNoneEvaluation, development, proof-of-concept

Pointing the chart at an external database is a small change, database.standalone: true with hostname, port, and credentials, and it converts the most fragile component of the deployment into someone else's operational problem. For a production RA-K1, it is the highest-value single decision available.

warning

Kasm writes platform logs to the database by default. At scale this reaches tens of gigabytes within days, and on Kubernetes it lands on a PersistentVolume that may not grow easily. Forward logs to an external SIEM and reduce internal retention before going to production. This is a readiness requirement, not a tuning suggestion.


Asymmetrical Failure Behavior​

RA2 introduces asymmetrical failure: an Agent failure is local, and a control plane failure is global. RA-K1 keeps that asymmetry and adds a second one along the orchestration axis.

Kasm Workspaces RA-K1 failure domains and blast radius diagram
FailureRecovered byBlast radius
A control plane pod crashesKubernetes, automaticallyContained to that replica. Invisible to users at medium or large.
A cluster worker node failsKubernetes, automaticallyPods reschedule. Only an in-cluster database PVC reattachment is slow.
The Agent VM failsNobody — an operator, or an AutoScale providerEvery container session on it dies. With one Agent, that is all of them. The UI stays up with nowhere to place new sessions.
Ingress controller or Primary Load Balancer failsKubernetes, only if replicatedTotal outage. Everything behind it is healthy and unreachable.
PostgreSQL failsNobody, if in-clusterTotal outage. Restart is not failover.
A remote RDP, VNC, or SSH target failsOutside Kasm entirelyContained to users of that target

The pattern is clear once stated: Kubernetes narrows the failure domains it owns to near-invisibility and leaves the ones it does not own completely untouched. Every remaining single point of failure in RA-K1 lives in a row Kubernetes cannot reach: the Agent VM, the ingress path, and the database.

info

RA-K1 is partially resilient, not highly available. A single Agent VM makes container sessions a single point of failure regardless of how many control plane replicas are running. It is appropriate for teams that want the control plane on Kubernetes, need session compute capacity beyond a single host, and can tolerate the loss of container sessions during Agent maintenance. It should not be presented as a production-grade HA architecture until the Agent tier has at least N+1 members and the database is externalized.


Capacity: Two Scaling Models in One Deployment​

RA-K1 has two capacity dimensions that scale by entirely different mechanisms, on entirely different timescales.

Control plane capacity scales the Kubernetes way. Change deploymentSize, or override components.<name>.replicas, and run helm upgrade. New pods are running in seconds. Capacity is bounded by cluster node resources, and adding nodes is a cluster operation independent of Kasm. This dimension governs API throughput, administrative UI responsiveness, and, importantly, RDP, VNC, and SSH session capacity, since kasm_guac does that work in-cluster.

Session capacity scales the virtual machine way. Provision a VM, install the Agent role against the manager hostname and token, then enable the Agent in the admin UI under Infrastructure → Docker Agents. This takes minutes, not seconds, and it involves a manual approval step that has no Kubernetes analogue. Capacity is bounded by the CPU and memory of the Agent VM pool.

Because the two are independent, the failure mode is predictable: an organization scales the thing it can see. The cluster is instrumented and has dashboards, and it responds to helm upgrade. The Agent VM is a machine somebody built once. Container session capacity quietly becomes the constraint while the control plane runs at ten percent utilization.

The mitigation is to treat the Agent pool as a first-class tier with its own capacity plan, its own monitoring, and, at production scale, its own AutoScaling configuration, so that the imperative side of the architecture acquires at least some of the elasticity the declarative side has by default.


Lifecycle: A Declarative Control Plane and an Imperative Agent​

Upgrade is where the two-orchestrator structure becomes most visible operationally, because the two halves upgrade by completely different procedures and must nonetheless stay version-compatible.

The control plane upgrades declaratively. Chart versions track Kasm versions: chart 1.1190.x corresponds to Kasm 1.19.0, and helm upgrade rolls the Deployments and StatefulSets with the chart's uptime policies respected. By default, the chart uses rolling image tags. Pinning to timestamped tags such as 1.19.0-rolling-20260215 makes builds immutable and upgrades explicit, which is the right choice for production.

The Agent upgrades imperatively. Someone connects to the VM, downloads the release, and runs the upgrade script. There is no desired-state reconciliation, no rollout status to watch, and no automatic rollback.

Two chart values govern the one-time database jobs, and they are the most common cause of a failed Kasm upgrade on Kubernetes:

dbManagement:
initialize: false # must be false after the first successful install
upgrade:
enable: false # true ONLY when the image change includes a schema migration

initialize: true left in place re-runs the initialization job on every helm upgrade and fails against an already-initialized database. upgrade.enable: true runs a full backup-and-restore cycle on every upgrade whether or not a migration is needed, which is safe but slow and unnecessary. Setting it to false when a migration was required produces pods that will not start and an alembic version mismatch in the api logs. The recovery is to re-run with upgrade.enable: true and then reset it.

tip

The practical discipline: keep both flags false in the committed values file, and flip upgrade.enable to true deliberately, for exactly one upgrade, when crossing a Kasm minor version.


Ports and Protocols​

Default ports. The Agent's listening port can be changed at install time with -L. If it is, the corresponding firewall and NetworkPolicy rules must change with it.

SourceDestinationPortPurpose
End userPrimary Load Balancer, to ingress443UI, API, authentication, session streaming
End userkasm_rdp_gateway Service3389Direct RDP for thick clients. Requires directRdpService.enabled: true with a NodePort or LoadBalancer Service.
kasm_api / kasm_manager podsPostgreSQL5432All platform state
kasm_manager podsAgent VM443Session provisioning and control instructions
Agent VMpublicAddr (through the Primary Load Balancer)443Check-in, registration, image requests, authentication
kasm_proxy podsAgent VM443KasmVNC session stream relay
kasm_proxy podsRemote systems running KasmVNC8443Proxied KasmVNC to VMs and hardware
kasm_guac podsRemote RDP systems3389Protocol translation to browser-native
kasm_guac podsRemote VNC systems5900Protocol translation to browser-native
kasm_guac podsRemote SSH systems22Protocol translation to browser-native
kasm_guac podskasm_api443Connection authorization and status checks
kasm_rdp_gateway podsRemote RDP systems3389Thick-client RDP passthrough

The line worth pausing on is the one that has no equivalent in any VM-based pattern: standard Kubernetes Ingress carries HTTP and HTTPS only. Every flow in this table on 443 fits that model naturally. Direct RDP on 3389 does not, and it needs its own Service. It is the single place where the Kubernetes networking model and the Kasm feature set do not line up cleanly, and it is worth deciding early whether the deployment needs it, because retrofitting an L4 path into an L7-only edge is more disruptive than provisioning it up front.


What RA-K1 Reveals About Kasm's Design Philosophy​

Role separation is real, not cosmetic. The chart deploys App, Dedicated Proxy, and Connection Proxy as independent workloads that scale and fail independently. These were always distinct roles. Kubernetes simply makes the distinction visible in a way that a multi-server VM install does not, because each role becomes a separately observable, separately scalable object.

The control plane is stateless, and this is what makes it portable. The App role can move to Kubernetes at all only because it holds no session state. Every replica is interchangeable, every request can land anywhere, and the entire tier can be replaced during a rolling upgrade without a user noticing. The property that RA3 relies on for horizontal scaling is the same property that lets RA-K1 exist.

The execution plane is deliberately not portable, and that is a design decision rather than a gap. Kasm could have shipped an Agent that runs in a pod. It does not, because the properties that make a good Kubernetes workload (disposable, rescheduleable, unprivileged, stateless) are the opposite of the properties an interactive desktop session needs. Recognizing which workloads belong on an orchestrator and which do not is the substance of the pattern.

The browser remains the universal client. Whether execution happens in a pod, in a container on a VM, or on a Windows server that predates the deployment by a decade, the user gets a WebSocket stream to a browser. RA-K1 spans three completely different execution substrates and presents one interface over all of them.


RA-K1 in the Deployment Progression​

RA-K1 does not replace a numbered pattern. It runs alongside them as the Kubernetes expression of the separation RA2 introduces. Its control plane resilience is closer to RA3, and its session plane resilience is closer to RA2 until the Agent tier is scaled out.

RA1RA2RA-K1RA3
Control plane locationOne VMOne VMKubernetes clusterMultiple VMs behind a load balancer
Control plane redundancyNoneNoneReplica count per componentN+1 Web App servers
Control plane recoveryManualManualAutomatic (Kubernetes)Load balancer health checks
Execution planeSame VMExternal Agent VMsExternal Agent VMsExternal Agent VMs
Database resilienceNoneNoneNone by default. Externalize it.Managed or replicated
ZonesOneOneOne (chart supports more)One
Failure modes an operator must handle by handAllAgent, control plane, databaseAgent, databaseDatabase

What RA-K1 lacks, and what each gap teaches:

What RA-K1 lacksWhat it teaches
A resilient database by defaultThat an orchestrator's self-healing does not extend to the one component that cannot simply be restarted
Agent tier redundancy in its minimal formThat session capacity and control plane capacity are independent dimensions that must be planned separately
Any reconciliation of the execution planeThat the boundary between orchestration domains is a real architectural boundary, not an implementation detail
Multi-zone session localityThat geographic distribution is a session plane problem, solved by adding zones and proxies, not by growing the cluster

RA-K1's value is that it makes a specific trade explicit. It takes the half of the platform that fits an orchestrator perfectly and hands it over completely, gaining declarative deployment, automatic recovery, and rolling upgrades for the control plane. It keeps the other half on infrastructure that models it correctly, accepting that this half must be operated the old way.

The mistake RA-K1 invites is assuming the first half's properties extend to the second. They do not, and the architecture is only sound when the Agent tier and the database are given the same deliberate resilience planning that Kubernetes provides for free everywhere else.


Appendix A: Minimum Production Values​

The production-shaped minimum, fully commented. For an evaluation deployment, delete the database.standalone block to use the bundled StatefulSet and drop deploymentSize to small.

# ---------------------------------------------------------------------------
# RA-K1 — Kasm Helm values for a Kubernetes control plane with an external
# Docker Agent VM in the same zone.
#
# Chart: oci://registry-1.docker.io/kasmweb/kasm-helm
# Version: 1.1190.x (chart 1.1190.x == Kasm 1.19.0)
# ---------------------------------------------------------------------------

# The single public hostname for the whole deployment. Users reach it, and the
# Agent VM reaches it. It must resolve and be reachable from BOTH.
publicAddr: kasm.example.com

# TLS. Either reference a pre-created secret in the namespace, or drive
# cert-manager with certificate.certManager.
certificate:
secretName: kasm-tls

# small = 1 replica/component, medium = 2, large = 3.
# NOTE: the database StatefulSet is ALWAYS 1 replica regardless of this value.
deploymentSize: medium

# ---------------------------------------------------------------------------
# Ingress
#
# The chart exposes the proxy in one of three mutually exclusive ways:
# 1. proxyService.type: LoadBalancer (chart default)
# 2. ingress.enabled: true (needs an ingress controller)
# 3. route.enabled: true (OpenShift)
#
# Whichever is chosen, replicate it. A single-replica ingress controller in
# front of a 3-replica control plane is still a single point of failure.
# ---------------------------------------------------------------------------
ingress:
enabled: true
tls: true
ingressClassName: nginx
annotations:
{}
# nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
# nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"

proxyService:
type: LoadBalancer

# ---------------------------------------------------------------------------
# Direct RDP for thick clients (port 3389).
#
# Standard Kubernetes Ingress carries HTTP/HTTPS only. Browser-based RDP works
# out of the box over 443 through the RDP HTTPS Gateway. Native RDP clients on
# 3389 need their own L4 Service. Decide this up front: retrofitting an L4
# path into an L7-only edge is disruptive.
# ---------------------------------------------------------------------------
directRdpService:
enabled: false
# type: LoadBalancer
# loadBalancerPort: 3389
# rdpAccessURL: rdp.example.com

# ---------------------------------------------------------------------------
# Database
#
# THE highest-value decision in RA-K1. The bundled kasm_db StatefulSet runs a
# single replica at every deploymentSize. Kubernetes restarts it; a restart is
# not a failover. Point at a managed service with automatic failover.
# ---------------------------------------------------------------------------
database:
standalone: true
hostname: kasm-db.example.internal
port: 5432
kasmDbName: kasm
kasmDbUser: kasmapp
# Credentials come from a secret you create in the namespace:
# kasmDbSecret:
# name: kasm-db-credentials
# key: password

# ---------------------------------------------------------------------------
# One-time database jobs
#
# Keep BOTH false in source control after the first successful install.
#
# initialize: true left in place re-runs the init job on every helm upgrade
# and fails against an already-initialized database.
#
# upgrade.enable Set to true deliberately, for exactly ONE upgrade, when
# crossing a Kasm minor version that carries a schema
# migration (e.g. 1.18.0 -> 1.19.0). Then reset to false.
# Getting this wrong the other way produces pods that will
# not start and an alembic version mismatch in the api logs.
# ---------------------------------------------------------------------------
dbManagement:
initialize: true
upgrade:
enable: false

# ---------------------------------------------------------------------------
# Image pinning
#
# The chart defaults to rolling tags, which are rebuilt on a schedule. That is
# convenient and non-deterministic. Pin to timestamped tags for production and
# advance them intentionally.
# ---------------------------------------------------------------------------
useImageTags: "1.19.0-rolling"

components:
api:
image:
tag: 1.19.0-rolling-20260215
manager:
image:
tag: 1.19.0-rolling-20260215
proxy:
image:
tag: 1.19.0-rolling-20260215
guac:
image:
tag: 1.19.0-rolling-20260215
rdpGateway:
image:
tag: 1.19.0-rolling-20260215
rdpHttpsGateway:
image:
tag: 1.19.0-rolling-20260215

imagePullPolicy: Always

# ---------------------------------------------------------------------------
# Scheduling
#
# Spread control plane replicas across nodes, otherwise "medium"/"large" buys
# replica count without buying node-failure tolerance.
# ---------------------------------------------------------------------------
affinity:
{}
# podAntiAffinity:
# preferredDuringSchedulingIgnoredDuringExecution:
# - weight: 100
# podAffinityTerm:
# topologyKey: kubernetes.io/hostname
# labelSelector:
# matchExpressions:
# - key: app.kubernetes.io/instance
# operator: In
# values: [kasm]

nodeSelector: {}

applySecurity: true
applyHealthChecks: true

Install:

kubectl create namespace kasm
helm install kasm oci://registry-1.docker.io/kasmweb/kasm-helm \
--version 1.1190.0 -n kasm -f my-values.yaml

Appendix B: Readiness Checklist​

Before calling an RA-K1 deployment production-ready:

  • Database is external and has automatic failover, not the bundled StatefulSet
  • Internal log retention reduced, and logs forwarded to a SIEM
  • Ingress controller replicated, or TLS terminated at a redundant external load balancer
  • At least two Agent VMs, so that Agent maintenance is not an outage
  • AGENT_HOSTNAME verified reachable from a pod, not just from a bastion: kubectl run -n kasm netcheck --rm -it --image=busybox --restart=Never -- nc -zv <AGENT_HOSTNAME> 443
  • MANAGER_HOSTNAME verified reachable from the Agent VM: curl -I https://kasm.example.com
  • Upstream Auth Address set to the load balancer address, if Agent AutoScaling is planned
  • dbManagement.initialize: false and upgrade.enable: false committed to the values file
  • Image tags pinned to timestamped builds
  • NetworkPolicy restricting namespace egress to the database and the Agent addresses only
  • Agent VM firewall permits 443 only from the cluster's egress addresses
  • Backup and restore of the database tested end to end, not merely configured

References​


This article is part of the Kasm Workspaces Reference Architecture Explanation Series.