COPY-PASTE HELPERS: @ @@@ | ~ {} [] \ ^ ` BOX: dist.gcp.emteria.com (permanent name; today 23.88.61.31, Hetzner bundle-transfer) JUMPHOST: browsers there go through the corporate proxy, which resolves *.idnow.vodafone.de publicly and lands on www.vodafone.de. Add *.idnow.vodafone.de to the proxy exceptions (or use no proxy) - confirmed 2026-09-29. Vodafone idnow, 2026-09-29: Cilium upgrade final12 -> final13, plus the real TLS certificate. kubectl on the node = /usr/local/bin/k3s kubectl (everything runs as root). firewalld is ON: the scripts run plain install.sh --upgrade. NTP WORKAROUND: the preflight dies with "clock is not NTP-synchronised" because chrony syncs from the hypervisor clock and systemd's NTPSynchronized flag stays "no". upgrade.sh now tries `chronyc makestep; chronyc waitsync`, and if the flag is still "no" it does the preflight by hand (SHA256SUMS, 2x bundle free under /var/lib, firewalld) and runs install.sh --upgrade --skip-preflight. Check the printed node clock against your laptop before answering y. Manual equivalent: cd /data/vodafone-idnow-final12 && sha256sum --quiet -c SHA256SUMS && df -h /var/lib && date -u && ./install.sh --upgrade --skip-preflight ("WARN: port 5000 already in use" on an upgrade is normal: that is the running bundle registry.) Every script echoes each step with a timestamp; upgrade.sh and tls-swap.sh ask y/N before changing anything. Setup once per terminal: export PROXY_ARG="-x http://172.16.13.110:3128" (only if the node needs the proxy) export B=/data ; cd /data 1. STEP A DOWNLOAD (done) curl -fsS $PROXY_ARG http://23.88.61.31/download.sh | bash -s -- final12 2. STEP A UPGRADE: Cilium 1.19.8 (before-state, backup, y/N, upgrade, verify, smoke) curl -fsS $PROXY_ARG http://23.88.61.31/upgrade.sh | bash -s -- final12 final11 expect: Cilium: Ok 1.19.8 | KubeProxyReplacement True | Masquerading BPF | Controller Status N/N healthy | 14/14 Synced Healthy 3. STEP B DOWNLOAD - 2nd terminal, same exports, while you look at step A curl -fsS $PROXY_ARG http://23.88.61.31/download.sh | bash -s -- final13 (laptop) merge gitops !163, close !159 4. TLS SWAP - independent of Cilium; run once step A is healthy certificate files are in /data: cert_mdm.idnow.vodafone.de.{crt,key,p7b,p7b.pem,pem} cd /data/vodafone-idnow-final12 curl -fsS $PROXY_ARG http://23.88.61.31/tls-swap.sh | bash -s -- /data/cert_mdm.idnow.vodafone.de.crt /data/cert_mdm.idnow.vodafone.de.key /data/cert_mdm.idnow.vodafone.de.p7b.pem (splits the p7b into leaf + CA chain, converts the DER .crt to PEM, seeds with a patched seed-secrets.sh, force-syncs the two ExternalSecrets) curl -fsS $PROXY_ARG http://23.88.61.31/tls-finish.sh | bash (restarts website + mdm, shows what 443/8883 serve, verifies against the Vodafone chain, waits for 14/14, smoke with the Vodafone CA bundle: dns.* and mdm.cert-match must PASS) curl -fsS $PROXY_ARG http://23.88.61.31/tls-verify.sh | bash (verification only, no restarts: served chain vs the cluster CA bundle + smoke with that CA) Facts found on the way: the delivered .crt is DER; the bundled seed-secrets.sh rejects the pair because it compares RSA moduli of a mixed DER+PEM file -> both handled by the scripts; repo fix follows. 5. STEP B UPGRADE: Cilium 1.20.2 cd /data curl -fsS $PROXY_ARG http://23.88.61.31/upgrade.sh | bash -s -- final13 final12 expect: Cilium: Ok 1.20.2 | Modules Health Degraded(0) | 14/14 Synced Healthy | smoke: dns.* PASS, mdm.cert-match PASS 6. OPERATOR TOOLS (installs /usr/local/bin/emteria-status and emteria-restart) curl -fsS $PROXY_ARG http://23.88.61.31/install-tools.sh | bash && source /etc/profile.d/emteria-path.sh emteria-status # ~10 s: k3s, node, cilium, 14/14 apps, pods, ESO, https api+hub, cert expiry, MQTTS, disks emteria-restart apps # rolling restart of the default-namespace workloads (DB + platform stay up) emteria-restart k3s # systemctl restart k3s (containers keep running) emteria-restart all # k3s-killall.sh + start: cold restart of everything (each restart level asks y/N, then waits for 14/14 and runs emteria-status) 7. SEND A FULL DIAGNOSTIC DUMP BACK (HTTPS PUT inbox on the box, Basic auth; Claude reads it over ssh) once, on the laptop, read the inbox password: ssh root@23.88.61.31 cat /root/inbox-credential NOTE: the node reaches port 443 only through the proxy. Maintenance runs as root (su -); export PROXY_ARG in that shell. export PROXY_ARG="-x http://172.16.13.110:3128" curl -fsS $PROXY_ARG http://23.88.61.31/report.sh | bash -s -- enroll # asks for that password once curl -fsS $PROXY_ARG http://23.88.61.31/report.sh | bash -s -- second 60 # dump with logs of the last 60 min (label free) (tar.gz with: df/mounts, what is on / vs /data (du, biggest files), containerd/registry/journal sizes, k3s, cilium, argo apps, all pods, ESO, PVCs, ingresses, CNPG, warning events, emteria-status, hub config; logs of ingress-nginx, the six services, mdm, postgres, argocd-repo-server, k3s journal; plus 01-errors-summary.txt with 4xx/5xx + 422/error lines) 8. FIX THE HUB 422s (found in the first dump, 2026-09-29 11:32Z) productmanager + storagemanager still ran the pods from install day. They mount customer-ca-bundle via subPath, which never picks up a Secret update, so their token call to https://api.idnow.vodafone.de failed with "certificate chain: UntrustedRoot" -> hub GET /product/v1/products and /storage/v1/files answered 422. tls-finish.sh only restarted website + mdm; it now restarts every deployment. One-off fix on the node: /usr/local/bin/k3s kubectl -n default rollout restart deployment /usr/local/bin/k3s kubectl -n default rollout status deployment/productmanager --timeout=300s emteria-status (then reload the hub: products + files pages must load) The two 422 on POST /main/api/v3/accounts/login were a wrong password, not a fault. 9. DISK: WHAT THE DUMP SAYS IS ON / AND ON /data / 24G, 15G used: /var/lib/rancher 12.1G (containerd image store 11.5G, k3s airgap tar 0.3G), kubelet 0.9G, /usr 1.6G. The three PersistentVolumes (postgres 5Gi, seaweedfs 20Gi, rabbitmq 2Gi) live in /var/lib/rancher/k3s/storage, i.e. ON ROOT (still tiny today) - uploaded device files will grow there. /data 246G, 19G used: emteria-registry 7.1G, vodafone-idnow-final12 5.9G, final13 5.9G. Tarballs already gone. Optional maintenance (stack down for a few minutes): move /var/lib/rancher onto /data with a bind mount curl -fsS $PROXY_ARG http://23.88.61.31/move-rancher.sh | bash (stops k3s, rsync to /data/rancher, fstab bind, restarts, waits for 14/14; keeps /var/lib/rancher.old-root until you rm it) ROLLBACK step A: /data/vodafone-idnow-final12/install.sh --rollback /data/vodafone-idnow-final11 final11 dir was deleted after install -> fetch it first (on the transfer box from ~10:45 UTC): curl -fsS $PROXY_ARG http://23.88.61.31/download.sh | bash -s -- final11 emergency without the dir (1.18.6 blobs are still in the node registry; upgrades only add): curl -fsS $PROXY_ARG -O http://23.88.61.31/final11-bootstrap.yaml /usr/local/bin/k3s kubectl apply --server-side -f final11-bootstrap.yaml step B: /data/vodafone-idnow-final13/install.sh --rollback /data/vodafone-idnow-final12 (never two minors back) TLS: cd /data/vodafone-idnow-final12 && ./seed-secrets.sh --self-signed idnow.vodafone.de --force, then curl -fsS $PROXY_ARG http://23.88.61.31/tls-finish.sh | bash (smoke will then need --ca /var/lib/emteria/selfsigned/tls.crt again) rollback does not undo a schema migration (none expected today). EXPECTED SMOKE NOISE (README section 10): one *-db-migrate pod Error WARN; argo.repo-server restarts equal to the before-state baseline; mdm app restart = the k3s restart / the TLS restart. dns.* rows must PASS now that DNS is set. AFTERWARDS /data/vodafone-idnow-final13/prune.sh /data # keeps 3 bundle dirs rm -f /data/bundle-vodafone-idnow-final1?.tar.zst laptop: README section 11 + shipped/ rows for final12/final13; delete Hetzner server + firewall bundle-transfer.