Storage
Longhorn is the only component whose failure mode is data loss, so it is the only one BKE refuses to touch when the cluster is not in a fit state.
Before it installs
components:
longhorn:
enabled: true
namespace: longhorn-system
secrets:
- longhorn-backup
settings:
backup:
target: s3://bke-backups@lon/prod
credentialSecret: longhorn-backup
defaultClassReplicaCount: 3
storageMinimalAvailablePercentage: 15
A cluster can run without a backup store. Set both backup.target and
backup.credentialSecret to "" and leave the secrets: list out; no Secret
and no /etc/bke/secrets.d/ directory are then needed. This suits an
evaluation cluster. For a cluster holding real data, configure a backup target
before the data exists: Longhorn is the one component whose failure mode is
data loss, and the backup-target refusal below can only protect backups that
happen.
defaultClassReplicaCount is worth deciding rather than defaulting. 3 is
right for production and needs at least three schedulable nodes; 1 is for a
single-node or test cluster and gives you no redundancy at all. Changing it later
affects new volumes, not existing ones.
dm_crypt must be available for encrypted volumes. BKE enables it at install
time on every node it provisions. A node provisioned before that existed, or one
whose kernel lacks the module, will attach ordinary volumes fine and fail to
attach encrypted ones.
Longhorn will use every schedulable node for replicas unless you tell it otherwise. If some nodes are not meant to carry storage, decide that before installing rather than after replicas have been placed.
The seven refusals
apply.sh checks the cluster’s storage state before the component loop, not
during it, so a refusal happens before anything has been changed. Each condition
has a name, and the name is what you type back if you decide to proceed anyway.
| Condition | Refuses when |
|---|---|
degraded-volume |
a volume is not fully replicated. Upgrading now risks the only healthy replica |
engine-skew |
volumes are running an engine image more than one minor behind the current one. One minor behind warns; beyond one refuses, because Longhorn supports the current engine and the previous one only |
backup-target |
the backup target is unavailable, or has not synced in over 24 hours |
incomplete-backup |
a backup is in progress |
node-not-ready |
a Longhorn node is not ready, so its replicas are not available |
minor-skip |
the upgrade would skip a Longhorn minor version |
downgrade |
the target is older than what is installed |
Each refusal names the condition and says what is wrong, on one line, so the name stays attached to its reason.
Proceeding anyway
BKE_ACCEPT_LONGHORN_RISK="backup-target" \
curl -sL https://bke.maml.uk/apply.sh | sh -s -- --version 2.1.0 --commit
It takes condition names, not a blanket yes. You can accept a stale backup target without also accepting a degraded volume. Accepting one is logged as a warning naming what was accepted.
There is no override for minor-skip or downgrade that makes them safe —
accepting them means you have decided the risk is yours.
Backups
Check your backup target actually works, once, when you first configure it:
kubectl -n longhorn-system get backuptargets.longhorn.io default \
-o custom-columns=URL:.spec.backupTargetURL,AVAILABLE:.status.available,SYNCED:.status.lastSyncedAt
AVAILABLE must be true and SYNCED must be recent. A target that has never
worked looks exactly like one that is idle, and the reason is only in the
longhorn-manager log:
kubectl -n longhorn-system logs -l app=longhorn-manager --tail=200 | grep -i backup
Two mistakes account for most failures:
A missing scheme on AWS_ENDPOINTS. s3.example.com is rejected as an
invalid URI; https://s3.example.com is accepted. The symptom is a target that
never becomes available.
A region that does not apply. If your object storage is not AWS, the region in the target URI is whatever that vendor uses, and copying one from an AWS example gives you a target that authenticates and then cannot find the bucket.
BKE checks the target’s availability and staleness as part of its preflight, so once it is configured you will hear about it going stale. It cannot check it before it exists.
Engine images do not move on their own
This is the one that surprises people.
Upgrading the Longhorn chart does not upgrade the engine image of existing volumes. Each volume pins the engine it was created with, and it keeps it. After a chart upgrade you can have a cluster running the new Longhorn with every volume still on the old engine — indefinitely.
The setting that moves them is Longhorn’s
concurrent-automatic-engine-upgrade-per-node-limit, and it defaults to 0,
meaning never.
This is a Longhorn setting, and BKE deliberately leaves it to you. Upgrading a volume’s engine detaches and reattaches it, and that is not something to trigger as a side effect of installing components — a cluster that is fine one moment should not lose volume attachments because someone applied a chart. BKE reports the skew on every run and refuses an upgrade that would outrun it; the moment to move the engines is yours to choose.
Our recommendation is 1:
kubectl -n longhorn-system patch settings.longhorn.io concurrent-automatic-engine-upgrade-per-node-limit --type=merge -p '{"value":"1"}'
One at a time per node, so volumes are interrupted gradually rather than together. A higher number is not faster in any way that helps you.
Because BKE does not manage this setting, your change is not overwritten by the
next apply.sh — and equally is not restored if something else removes it. Check
it after any Longhorn change:
kubectl -n longhorn-system get settings.longhorn.io concurrent-automatic-engine-upgrade-per-node-limit -o jsonpath='{.value}'
check.sh reports this value, and says plainly when it is 0 that Longhorn will
not upgrade engines on its own.
Check what your volumes are actually on:
kubectl -n longhorn-system get volumes.longhorn.io \
-o custom-columns=NAME:.metadata.name,ENGINE:.status.currentImage,STATE:.status.state
More than one version in that column is skew. One minor behind is reported as outstanding work; more than one minor behind refuses the next upgrade, because Longhorn itself supports only the current engine and the one before it.
The reporting starts at the first minor deliberately. Leaving this setting at 0
is exactly how a cluster reaches three minors of skew without anyone noticing,
and by then the fix is large. Told at one minor, it is small.
Before a cluster upgrade
Longhorn is the reason to take the ordering in Upgrading seriously:
- Confirm the backup target is available and recently synced.
- Confirm no volume is degraded —
kubectl -n longhorn-system get volumes.longhorn.ioand look for anything notattached/healthy. - Confirm every volume is on the same engine image.
- Then start.
All four are what the preflight checks. Checking them yourself beforehand means that a refusal, if it comes, arrives before the maintenance window rather than during it.