BKE Bahriya Kubernetes Engine Start a 14-day trial

Adding nodes

Every additional node is two steps: BKE provisions it, then kubeadm joins it. The join is yours to run. Whether a node should be a control-plane member or a worker — and why control-plane nodes come in odd numbers — is covered in How a cluster fits together.

Each section below names the machine its commands run on. Two machines are involved every time: the new node being added, and the first control-plane node, where join credentials are minted.

BKE does not automate the join deliberately. Joining needs a token and a CA hash that are minted on the first control-plane node and expire, and a join token is a credential that admits a machine to your cluster — it is not something an installer should be fetching over the network on your behalf.

Run check.sh on each node before provisioning it.

Provision the node

On the new node:

curl -sL https://bke.maml.uk/install.sh | sh -s -- \
  --role worker \
  --version 2.0.0 \
  --hostname w-bukhara.prod.k8s.example.com

Use the same --version as the rest of the cluster. --endpoint and --dns-domain are not passed: they belong to the cluster, which already exists, and this node will learn them when it joins.

This does everything the first node’s provisioning did except kubeadm init — packages, kernel modules, containerd, UFW, swap off, sysctls. When it finishes, the node is ready to join and has not joined.

Use --role worker for additional control-plane nodes too. The role decides whether the script initialises a cluster, and only the very first node does that. What makes a node a control-plane member is the join flag below.

Join a worker

On the first control-plane node:

kubeadm token create --print-join-command

Run what it prints on the new node, as root. It looks like:

kubeadm join prod.bke.example.com:6443 \
  --token <token> \
  --discovery-token-ca-cert-hash sha256:<hash>

Tokens expire after 24 hours by default. Mint a fresh one rather than reusing an old one you have written down.

Join a control-plane node

Control-plane members need one more thing: the certificate key that encrypts the shared certificates uploaded by --upload-certs during kubeadm init. It expires after two hours, so mint it when you are ready to join.

On the first control-plane node:

kubeadm init phase upload-certs --upload-certs

That prints a certificate key. Then take a join command as above and add two flags:

kubeadm join prod.bke.example.com:6443 \
  --token <token> \
  --discovery-token-ca-cert-hash sha256:<hash> \
  --control-plane \
  --certificate-key <certificate-key>

Add control-plane nodes in odd numbers. etcd needs a majority to accept writes, so three members tolerate one failure and four still tolerate only one. Two is worse than one: either failure takes the cluster down.

Your --endpoint must already resolve to something that can reach all of them before you add the second one. If it points at the first node’s own address, fix that first — it is not changeable afterwards.

kubectl and k9s on the new control-plane node

Provisioning installs kubectl and k9s on every node, but neither works on an additional control-plane member until it has a kubeconfig: the join wrote the administrator credential to /etc/kubernetes/admin.conf, and both tools read ~/.kube/config. Copy it once, as root, on the new control-plane node:

mkdir -p ~/.kube
cp /etc/kubernetes/admin.conf ~/.kube/config
chmod 600 ~/.kube/config

kubectl get nodes and k9s then work from this node exactly as they do from the first one.

Workers get no administrator credential from their join, and that is correct — a worker holds only its own kubelet’s identity. Run kubectl and k9s from a control-plane node rather than copying admin.conf onto machines that do not need it; Cluster login covers access for people who are not holding root on a control-plane node.

Check

On the first control-plane node:

kubectl get nodes -o wide

Every node should reach Ready. A node stuck in NotReady is almost always one of two things, in this order:

  1. Node-to-node ports are blocked somewhere BKE cannot see — a cloud security group, a firewall between subnets. UFW on the node allows the traffic; something in front of it may not.
  2. Calico has not finished on the new node. kubectl get pods -n calico-system -o wide shows whether its pods are running there.

Labelling

BKE does not label or taint nodes for you. Roles beyond control-plane — storage nodes, ingress nodes — are your convention, applied with kubectl label.

One case is worth knowing before you install components: Longhorn will use every schedulable node for replicas unless you tell it otherwise. If you intend some nodes to carry storage and others not, decide that before Longhorn is installed rather than after it has placed replicas.