config.yaml
One file at /etc/bke/config.yaml describes your cluster: what it is called,
which components you want, and how each is configured.
apply.sh sends it to api.maml.uk, which checks it and returns what to
install. An error in it comes back as a sentence naming what is wrong, before
anything on your cluster is touched.
This file leaves your cluster; /etc/bke/secrets.d/ never does. That is why
credentials live in a separate directory rather than in here.
Two files leave your cluster, and this is one of them. The other is
/etc/bke/values/<component>.yaml — see Installing components. Neither is
stored. secrets.d/ is sent nowhere, ever.
It describes the cluster. It does not say which BKE version you are running —
that is --version on whichever script you are running, because it is what you
are doing today rather than a property of the cluster.
The shape
schemaVersion: 1
cluster:
name: prod
region: lon
podNetworkCidr: 192.168.0.0/16
secrets:
unmanaged:
- longhorn-backup
auth:
enabled: false
components:
metrics-server:
enabled: true
namespace: kube-system
# ...
A complete file to start from
The file below is a full config.yaml for a new cluster, with every component
enabled and each in its conventional namespace. Copy it to
/etc/bke/config.yaml and replace every value written as <...> with your
own. The file is ready when
grep '<' /etc/bke/config.yaml
finds nothing. A placeholder you miss does not install something wrong: a value
config.yaml must set and does not is refused before your cluster is touched.
schemaVersion: 1
cluster:
# The name appears in your logs, your Longhorn backup paths and your netdata
# hostnames. Choose one that will still be recognisable in an alert.
name: <cluster-name>
region: <region>
# For a new cluster this is a choice, made here, before the first node is
# installed. 192.168.0.0/16 is a reasonable one unless it overlaps a network
# your nodes must reach.
podNetworkCidr: 192.168.0.0/16
auth:
# Read *Cluster login* before setting this to true.
enabled: false
reporting:
# A daily version report to us. Off is supported and costs nothing - see
# the reporting section on this page.
enabled: true
components:
metrics-server:
enabled: true
namespace: kube-system
cert-manager:
enabled: true
namespace: cert-manager
settings:
leaderElectionNamespace: kube-system
longhorn:
enabled: true
namespace: longhorn-system
# BKE creates this Secret from /etc/bke/secrets.d/longhorn-backup/ -
# create that directory with the target's credentials in it first.
# See *Secrets* for the directory's shape and *Storage* for the keys.
# To evaluate without a backup store, set both backup values to ""
# and remove this secrets list; no Secret is then needed.
secrets:
- longhorn-backup
settings:
backup:
target: s3://<bucket>@<region>/<path>
credentialSecret: longhorn-backup
# Replicas per volume: 1 on a single node, 3 in production.
defaultClassReplicaCount: 3
storageMinimalAvailablePercentage: 15
kong:
enabled: true
namespace: kong
kuma:
enabled: true
namespace: kuma-system
fluent-bit:
enabled: true
namespace: logging
settings:
output:
# A complete [OUTPUT] section, verbatim, written flush left within the
# block. *Logging with Fluent Bit* carries worked examples for
# OpenSearch, Loki and S3-compatible stores.
config: |
[OUTPUT]
Name loki
Match *
Host <log-destination-host>
Port 3100
netdata:
enabled: true
namespace: netdata
settings:
cloud:
# As written, this cluster is not claimed into Netdata Cloud and the
# named Secret need not exist. To claim it: set enabled to true, set
# room to your room UUID, put the claim token under
# /etc/bke/secrets.d/netdata-cloud/, and list netdata-cloud under this
# component's secrets - see *Secrets*.
enabled: false
room: ""
secret: netdata-cloud
Every key in it is explained below. Two habits worth keeping from the first
day: a component you do not want is enabled: false rather than deleted, so
the file still names it the day you change your mind; and nothing that is a
credential goes in this file — /etc/bke/secrets.d/ exists for those and
never leaves your cluster.
Top level
| Key | Required | Notes |
|---|---|---|
schemaVersion |
yes | Must be 1 |
cluster.name |
yes | Appears in your logs, your Longhorn backup paths and your netdata hostnames. Choose a name that will still be recognisable in an alert |
cluster.region |
yes | A label you choose. BKE does not interpret it |
cluster.podNetworkCidr |
yes | Must match what the cluster was created with. Changing it here does not change the cluster |
secrets.unmanaged
Names Secrets that your cluster already owns and BKE must not create. Use it when a Secret is produced by sealed-secrets, the External Secrets Operator, or by hand.
Every name listed here must also appear in some component’s secrets:
list. That declaration is what tells BKE which component needs the Secret —
and therefore which namespace to look for it in. secrets.unmanaged then
decides who provides it: for a name listed here, BKE checks the Secret exists
and creates nothing; for a declared name not listed here, BKE creates it from
/etc/bke/secrets.d/.
A name here that no component declares is refused — it is almost always a
spelling that no longer matches the secrets: list it was meant to pair with.
A Secret that is only referenced in settings — a credentialSecret value,
netdata’s cloud.secret — is not a declaration, and BKE neither creates nor
checks it.
A declared name that resolves to neither a directory under
/etc/bke/secrets.d/ nor an unmanaged entry is a pre-flight refusal —
before the first chart is fetched, not part-way through. See Secrets.
auth
auth:
enabled: false
oidc:
- issuerUrl: https://login.example.com/v2.0
clientId: bke-prod
usernameClaim: email
groupsClaim: groups
groupsPrefix: "oidc:"
rbac:
clusterAdminGroups: [oidc:platform-admins]
adminGroups: [oidc:team-leads]
editGroups: [oidc:developers]
viewGroups: []
enabled defaults to false, and the rbac groups can be listed while it is
off — the bindings are created and simply bind groups nobody can present yet.
That pairing is deliberate: on an existing cluster, enabling authentication
restarts your API server, and it should be a separate decision from installing
components.
enabled decides how your cluster is created, so on a new cluster it has to
be right before install.sh --role master runs. With it off, the cluster is
created exactly as it always was. usernamePrefix is optional and does for
usernames what groupsPrefix does for groups.
Read Cluster login before setting enabled: true. There is a limitation
there that will decide whether your identity provider can be used at all, and a
break-glass procedure to carry out first rather than afterwards.
reporting
reporting:
enabled: true
enabled defaults to true and is the only key. Once a day, a CronJob named
bke-report in the bahriya-system namespace sends us a short report about
this cluster. It is the one recurring thing that leaves your cluster, which is
why it has a first-class off switch.
Setting enabled: false removes the job, along with its ServiceAccount and
all of its RBAC, on the next apply.sh --commit. It is not merely no longer
installed: a cluster that reported yesterday stops reporting today.
Turning it off is a supported choice, not a degraded one. Nothing expires for silence and nothing is withheld.
Cluster reporting is the whole of it: every field that is sent, what the job’s permissions do and do not allow it to read, how to run it by hand and read the result on your own cluster, and what to check after turning it off.
components
Each component takes enabled, namespace, optionally releaseName, optionally
secrets, and optionally settings.
components:
longhorn:
enabled: true
namespace: longhorn-system
secrets:
- longhorn-backup
settings:
backup:
target: s3://bke-backups@lon/prod
credentialSecret: longhorn-backup
defaultClassReplicaCount: 3
storageMinimalAvailablePercentage: 15
namespace is required for every enabled component and has no default. That
is not an oversight. Four of the seven components are conventionally installed in
a namespace that is not the component’s own name — kube-system,
longhorn-system, kuma-system, logging — so a default would be wrong more
often than right, and being wrong means quietly installing a second copy
somewhere you were not looking.
releaseName defaults to bke-<component> and exists for clusters that came
before BKE. A new cluster gets bke-longhorn, bke-kong and so on, which is
what makes BKE’s releases obvious in helm list without a selector.
A cluster you already run almost certainly calls them something else — plain
longhorn, plain kong, or a name of your own. Name it here and BKE takes that
release over in place:
components:
kong:
enabled: true
namespace: kong
releaseName: kong
Renaming a Helm release means deleting and recreating it, which on Longhorn
or an ingress gateway is an outage, so BKE never does it. The name is a
convenience; what marks a release as BKE’s is the
app.kubernetes.io/managed-by=bke label, which is neither optional nor
configurable.
This is also why BKE takes over the release you name rather than one it
recognises. A release already called kong that BKE did not install is adopted
only because you asked for Kong in this file — never because something scanned
your cluster and found a familiar name.
Naming a component this BKE version does not ship is refused, not skipped. Asking for something this BKE version does not describe and receiving everything else would leave you believing you had it.
Per-component settings
| Component | Key | Meaning |
|---|---|---|
| cert-manager | leaderElectionNamespace |
Where cert-manager takes its leader-election lease. Conventionally kube-system |
| longhorn | backup.target |
The backup target URI, e.g. s3://bucket@region/path. Set it to "" to run without backups |
| longhorn | backup.credentialSecret |
Name of the Secret holding the target’s credentials. Set it to "" when backup.target is "" |
| longhorn | defaultClassReplicaCount |
Replicas per volume in the default StorageClass. 1 on a single node; 3 in production |
| longhorn | storageMinimalAvailablePercentage |
Longhorn stops scheduling to a disk below this |
| fluent-bit | output.config |
A complete [OUTPUT] section, verbatim. Write it flush left; it is re-indented for you. Logging with Fluent Bit carries worked examples |
| netdata | cloud.enabled |
Whether to claim this cluster into Netdata Cloud |
| netdata | cloud.room |
The Netdata Cloud room UUID to claim into |
| netdata | cloud.secret |
Name of the Secret holding the claim token |
Running without backups is written down, never inferred. To install
Longhorn with no backup store — reasonable while evaluating, not for a cluster
holding data you cannot recreate — set both backup.target and
backup.credentialSecret to "", leave secrets: out of the longhorn block,
and create nothing under /etc/bke/secrets.d/. Omitting the backup keys
entirely is refused rather than read as this choice, so a forgotten backup
block cannot quietly become a cluster that does not back up.
secrets: and settings: are two different statements. secrets declares
that the component needs a Secret by that name — it is the list BKE works from,
with secrets.unmanaged deciding per name whether BKE creates it from
/etc/bke/secrets.d/ or verifies it and touches nothing.
settings.*.credentialSecret and its like are references inside the
component’s configuration: they tell the component which Secret to use, and BKE
does not read them as declarations.
So for a Secret produced by something else — sealed-secrets, the External
Secrets Operator, your own kubectl — name it in settings, declare it in
secrets, and list it in secrets.unmanaged. BKE then checks it exists before
anything is applied and never creates, changes or deletes it. Declaring it
without the unmanaged entry tells BKE to manage it: its contents become
whatever /etc/bke/secrets.d/<name>/ holds, and every workload referencing it
restarts when that changes — the opposite of what you meant for a Secret
something else owns.
What an error looks like
A sentence in plain text, naming what is wrong and what to do about it.
It arrives before your cluster is touched. Everything in config.yaml is
checked while apply.sh is still gathering what it needs, so a mistake here
costs you a re-run and nothing else. BKE would rather stop than install
something half-described.