BKE Bahriya Kubernetes Engine Start a 14-day trial

Before you begin

BKE turns a set of Debian servers into a Kubernetes cluster. It installs upstream software — kubelet, containerd, Calico, Helm — pinned to versions we have tested together, and then installs the components you have chosen on top.

This page is the one worth reading in full before you touch a machine. Most of what goes wrong in an installation was decided before it started.

If control plane, worker and kubectl are new words, read How a cluster fits together first — every page from here on assumes them, and it also carries the table of which command runs on which machine.

What a machine needs

Debian 12 (bookworm) or Debian 13 (trixie). Both are tested against every release before it is published, and they are the only two this version installs on.

Another Debian release is refused. BKE installs the container runtime from a specific package that each BKE version pins for a named Debian release, and this version names only bookworm and trixie. There is no fallback: installing another release’s package would put it on a machine that did not ask for it. check.sh reports this, and install.sh stops before it changes anything on the machine.

Earlier BKE versions took the runtime from a third-party apt repository, which is why an untested Debian used to be a warning rather than a refusal. That repository also meant the same BKE version could install a different runtime depending on the day it ran. The pin removes both.

Anything that is not Debian is refused for the same reason, and is caught earlier still — before the version graph is even fetched.

Kubernetes packages come from pkgs.k8s.io, which is the only third party a node reaches.

Root. Every script is run as root or under sudo. They set hostnames, write to /etc/, manage apt, and load kernel modules.

amd64 or arm64. Helm is fetched for the architecture the node reports. Anything else is refused rather than guessed at.

Disk, in two places that fill at different rates:

Path Refuses below Comfortable
/ 5 GB 10 GB
/var/lib 20 GB 40 GB

/var/lib is the one to size generously. Every container image the cluster ever pulls lands there, and it grows with what you run rather than with what you installed.

curl on PATH. Nothing can be fetched without it. Everything else BKE needs — jq, kubectl, tar, etcd-client — is installed for you.

Swap will be turned off. install.sh runs swapoff -a and comments the swap entries out of /etc/fstab. This is not a preference; the kubelet refuses to start with swap enabled. If a machine is running something else that expects swap, it is not a candidate for a cluster node.

A clock within 60 seconds of correct. Certificate validation tolerates some drift; kubeadm tolerates about five minutes. Beyond a minute you are relying on that tolerance, and the failure — a node that joins and is then rejected — does not name the clock as the cause. Run systemd-timesyncd, chrony or ntpd before anything else.

What BKE reaches over the network

Every node, control-plane and worker alike, needs outbound HTTPS to all of these. There is no offline installation, and no bundle that removes the requirement.

Host What comes from it
bke.maml.uk the scripts themselves, and the version graph
mirrors.bke.maml.uk Calico manifests, Helm, the containerd package, calicoctl and k9s
charts.bke.maml.uk the Helm chart for every component
api.maml.uk your licence, and the plan and values for your cluster
pkgs.k8s.io kubelet, kubeadm, kubectl
prod-cdn.packages.k8s.io where pkgs.k8s.io redirects the package downloads
1x.ax every container image
your Debian mirror base packages

The two k8s.io names travel together: the Kubernetes project’s repository answers the package request with a redirect to its content delivery network, and apt follows it. A rule that allows only pkgs.k8s.io passes every check and then fails at the download step of an install or upgrade.

Every container image BKE installs comes from 1x.ax — the Kubernetes control plane, all seven components, and the pod sandbox image the container runtime pulls on its own account. They are copies of the upstream images, served from one host so that the rule you write does not change when a component does.

This is why the list above is short. Pulled from their own publishers, those images would need four registries and, because a registry answers a pull from more than one host, five further names for token services and content delivery networks — one of which is chosen by where the node is, so it cannot be written down in advance.

Your own workloads are unaffected. Nothing here changes how an image you deploy is resolved: a Pod that asks for nginx:latest still pulls it from Docker Hub. Whatever registries your own applications use remain your firewall rule to write, exactly as they were.

github.com, get.helm.sh and download.docker.com are no longer reached at all. Helm, the containerd package, calicoctl and k9s are served from mirrors.bke.maml.uk alongside the Calico manifests. Every one of them is a version a BKE release pins, and a pin whose bytes a third party can move or withdraw is not a pin — mirroring them is what makes the pin true, and it shortens the list above at the same time.

k9s and calicoctl remain operator conveniences rather than part of the cluster: if the mirror is unreachable when they are fetched, provisioning finishes and prints the command to install them later.

If your estate reaches the internet through a proxy, these are the hosts to allow before you start rather than after the first failure. check.sh probes every one of them and names any that did not answer.

What BKE opens between nodes

install.sh configures UFW on every node it provisions:

Port For
6443 the Kubernetes API server
2379, 2380 etcd (client and peer)
10250, 10257, 10259 kubelet, controller-manager, scheduler
179 BGP, for Calico
4789 VXLAN
5473 Calico Typha
51820, 51821 WireGuard, for encrypted pod traffic
22 SSH, or whichever port your sshd_config already names

BKE does not change your SSH configuration

/etc/ssh/sshd_config is yours. install.sh does not rewrite it, does not move the port, does not change your authentication policy, and does not restart sshd. Nothing about running Kubernetes requires any of that.

UFW opens the port your sshd_config already names, read from the file rather than assumed, so enabling the firewall cannot lock you out of a machine you are connected to. If the file names several ports, all of them are opened.

If you would rather BKE owned it, install.sh --manage-ssh installs our configuration: password authentication disabled, root permitted to log in by key only, and --ssh-port to choose the port. It applies on that node and on that run, and it replaces the file rather than merging into it. Before using it, confirm that key-based login already works and that the port you choose is open on anything in front of the machine.

Opening them in UFW is not the same as them being reachable. If the machines sit behind a cloud security group, a hardware firewall or separate subnets, those also have to allow this traffic, and nothing in BKE can see that they do. This is the single most common cause of a first node that comes up fine and a second node that will not join.

Three decisions you cannot change afterwards

These are arguments to install.sh, and all three are baked into the cluster the moment the first control-plane node comes up. Changing any of them later is a cluster rebuild, not an edit.

The node hostname becomes the machine’s hostname and its --node-name in Kubernetes, and it appears in the node’s certificates.

The control-plane endpoint — host:port, no https:// — is the stable address of the API server. It goes into every kubeconfig ever generated for this cluster, and it is the address your licence is issued against. A licence covers one cluster, and this is how we know which. Reference architecture describes what this address should point at in production — a load balancer across the control-plane nodes — and is worth reading before you choose it.

The service DNS suffix is what your services resolve under, as <service>.<namespace>.svc.<this>. cluster.local is the Kubernetes convention. There is no default: the installer requires you to state it.

Pick a naming convention now rather than for the first cluster only. A shape that works well is <role>.<cluster>.k8s.<your-domain> for node hostnames — m1.<cluster> for the first control-plane node, w-<city>.<cluster> for workers — and a separate <cluster>.bke.<your-domain> for the endpoint, so the address of the control plane is not tied to any one machine’s name. None of this is enforced; it is simply the shape that stops the second cluster from being awkward.

Then check, before you install

check.sh runs on every node, before an install and before an upgrade. It reads only — no package is installed, nothing is applied to a cluster, nothing is left behind. It is safe to run at any time, including in the middle of an upgrade.

curl -sL https://bke.maml.uk/check.sh | sh -s -- --version 2.0.0

It verifies every host above from this machine, checks the disk floors, the clock, swap and the tools, confirms your licence will not refuse later, and — on a control-plane node of an existing cluster — checks that the cluster is fit to be changed.

Exit Meaning
0 ready
1 a check failed; the output names which and what to do
2 the preflight could not run at all

The difference between 1 and 2 is deliberate. A failed check is information about your node. Exit 2 means the check itself could not reach a conclusion, and a preflight that cannot run has told you nothing — it must not be mistaken for one that passed.

Skipped checks are counted and printed for the same reason. If a check does not apply to this machine, you should be able to see that it did not apply rather than assume it passed.

What check.sh does not tell you

It runs on one machine and reports on that machine. It cannot see:

  • whether your nodes can reach each other on the ports above — it tests outbound reachability to the internet, not node-to-node connectivity;
  • whether your load balancer or DNS points at the endpoint you are about to declare;
  • whether the hardware is adequate for what you intend to run, as opposed to adequate to install.

Run it on every node, not only the first. A node that cannot reach what it needs will fail part-way through its own installation.