Files
nomad-csi/csi/controller.nomad
Henrik Jess Nielsen 53d696ea73
All checks were successful
Deploy CSI Jobs / deploy (push) Successful in 36s
fix(csi): survive int reboots, pin image by digest, raise memory
Both plugin jobs are pinned to int with count = 1 and had no restart or
reschedule stanza. When int goes away the allocations are marked Lost and
never return — csi-nfs-node showed 1 Complete / 3 Lost / 0 Failed, so these
were never application crashes, they were the node disappearing. The plugin
has been down since 2026-06-08 as a result. Adds restart + unlimited
reschedule with exponential backoff so they recover on their own.

The image was :latest, against our own rule for Nomad. The copy cached on
int is sha256:944d0e65 and roughly seven months old, so any fresh pull on
another node would get a different build — a good way to make failures
irreproducible. Pinned to the digest that has actually been running.

Memory was 128 MB for a Node.js process. Raised to 256 with memory_max so
mount and provisioning activity can burst without permanently reserving it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 23:22:45 +02:00

90 lines
2.2 KiB
HCL

job "csi-nfs-controller" {
datacenters = ["dc1"]
type = "service"
namespace = "default"
# Pin controller to int — NFS server lives here
constraint {
attribute = "${node.unique.name}"
value = "int"
}
group "controller" {
count = 1
# Pinned to int, so an int reboot takes the controller down with it.
# Without these the allocation is marked Lost and never returns.
restart {
attempts = 5
interval = "10m"
delay = "30s"
mode = "delay"
}
reschedule {
unlimited = true
delay = "30s"
delay_function = "exponential"
max_delay = "5m"
}
task "plugin" {
driver = "docker"
config {
# Pinned by digest rather than :latest — see node.nomad for why.
image = "democraticcsi/democratic-csi@sha256:944d0e65077efbd9c1fdf23997eec8fac4b4bfb7c3de400f63e33a0a849c5ced"
command = "/bin/democratic-csi"
args = [
"--csi-version=1.5.0",
"--csi-name=org.democratic-csi.nfs",
"--driver-config-file=${NOMAD_TASK_DIR}/driver-config-file.yaml",
"--log-level=info",
"--csi-mode=controller",
"--server-socket=${CSI_ENDPOINT}",
]
volumes = [
"/opt/csi-volumes:/opt/csi-volumes",
]
privileged = true
}
csi_plugin {
id = "org.democratic-csi.nfs"
type = "controller"
mount_dir = "/csi"
}
template {
data = <<TMPL
driver: nfs-client
instance_id: int-nfs-1
nfs:
shareHost: 192.168.15.25
shareBasePath: /opt/csi-volumes
controllerBasePath: /opt/csi-volumes
shareAlldirs: false
shareAllowedNetworks:
- 192.168.15.0/24
shareAllowedHosts: []
shareMaprootUser: root
shareMaprootGroup: root
server:
port: 50051
TMPL
destination = "${NOMAD_TASK_DIR}/driver-config-file.yaml"
}
# The controller does the provisioning work (volume create, NFS export
# management), so it gets more headroom than the node plugin.
resources {
cpu = 100
memory = 256
memory_max = 768
}
}
}
}