Files
nomad-csi/csi/node.nomad
Henrik Jess Nielsen 53d696ea73
All checks were successful
Deploy CSI Jobs / deploy (push) Successful in 36s
fix(csi): survive int reboots, pin image by digest, raise memory
Both plugin jobs are pinned to int with count = 1 and had no restart or
reschedule stanza. When int goes away the allocations are marked Lost and
never return — csi-nfs-node showed 1 Complete / 3 Lost / 0 Failed, so these
were never application crashes, they were the node disappearing. The plugin
has been down since 2026-06-08 as a result. Adds restart + unlimited
reschedule with exponential backoff so they recover on their own.

The image was :latest, against our own rule for Nomad. The copy cached on
int is sha256:944d0e65 and roughly seven months old, so any fresh pull on
another node would get a different build — a good way to make failures
irreproducible. Pinned to the digest that has actually been running.

Memory was 128 MB for a Node.js process. Raised to 256 with memory_max so
mount and provisioning activity can burst without permanently reserving it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 23:22:45 +02:00

95 lines
2.5 KiB
HCL

job "csi-nfs-node" {
datacenters = ["dc1"]
type = "service"
namespace = "default"
# Only int for now — expand to all workers once verified
constraint {
attribute = "${node.unique.name}"
value = "int"
}
group "node" {
count = 1
# int is both the NFS server and the only node running this plugin, so a
# reboot there takes it down. Without these the allocation is marked Lost
# and never comes back on its own — which is how it stayed dead from
# 2026-06-08 until someone noticed.
restart {
attempts = 5
interval = "10m"
delay = "30s"
mode = "delay"
}
reschedule {
unlimited = true
delay = "30s"
delay_function = "exponential"
max_delay = "5m"
}
task "plugin" {
driver = "docker"
config {
# Pinned by digest rather than :latest. The tag moves upstream, so a
# fresh pull on a new node gets a different build than the one cached
# on int — this digest is what has actually been running there.
image = "democraticcsi/democratic-csi@sha256:944d0e65077efbd9c1fdf23997eec8fac4b4bfb7c3de400f63e33a0a849c5ced"
command = "/bin/democratic-csi"
network_mode = "host"
args = [
"--csi-version=1.5.0",
"--csi-name=org.democratic-csi.nfs",
"--driver-config-file=${NOMAD_TASK_DIR}/driver-config-file.yaml",
"--log-level=info",
"--csi-mode=node",
"--server-socket=${CSI_ENDPOINT}",
]
privileged = true
}
csi_plugin {
id = "org.democratic-csi.nfs"
type = "node"
mount_dir = "/csi"
}
template {
data = <<TMPL
driver: nfs-client
instance_id: int-nfs-1
nfs:
shareHost: 192.168.15.25
shareBasePath: /opt/csi-volumes
controllerBasePath: /opt/csi-volumes
shareAlldirs: false
shareAllowedNetworks:
- 192.168.15.0/24
shareAllowedHosts: []
shareMaprootUser: root
shareMaprootGroup: root
mountOptions:
- nolock
- nfsvers=4
server:
port: 50051
TMPL
destination = "${NOMAD_TASK_DIR}/driver-config-file.yaml"
}
# democratic-csi is a Node.js process; 128 MB is tight enough that mount
# activity can push it over. memory_max lets it burst without reserving
# the headroom permanently.
resources {
cpu = 100
memory = 256
memory_max = 512
}
}
}
}