Kubernetes v1.37: Hardening Container Storage with Bind Mount Options and EmptyDir Permissions

Kubernetes v1.37 brings important storage security features: emptyDir permission modes and bind mount options. They help application programmers and security professionals implement rigorous security policies, for example, prohibiting deletion of files across containers or execution of arbitrary binaries from writable volumes, directly in Kubernetes without any complicated circumvention.

Linux storage and permission fundamentals

Before diving into the new Kubernetes features, let us briefly review the low-level Linux security mechanisms that make them possible.

Bind mount flags

When Linux mounts or remounts a directory, Virtual File System (VFS) flags control what actions are permitted on that filesystem:

  • noexec: Do not permit direct execution of any binaries on the mounted filesystem.
  • nosuid: Do not allow set-user-identifier or set-group-identifier bits to take effect.
  • nodev: Do not interpret character or block special devices on the file system.

Directory permissions and the sticky bit

Standard Unix permissions regulate access across three scopes: Owner, Group, and Others (e.g., 0755 or 0777).

Beyond standard read, write, and execute bits, Linux supports the sticky bit (as in mode 01777). When applied to a directory, the sticky bit ensures that a file inside that directory can only be deleted or renamed by the file's owner or root. This is essential for shared writable directories like /tmp.

Motivation for the improvements

Why does Kubernetes need bind mount options and emptyDir permissions?

The primary goal of these features is to increase the security of Kubernetes workloads by allowing security-related bind mount options on volume mounts. By default, volumes are bind-mounted into containers by the container runtime and kubelet without noexec, nosuid, or nodev flags. This default can undermine security. For example, with noexec missing, a compromised process can use any writable volume (emptyDir, PersistentVolume, etc.) to download, chmod +x, and execute arbitrary binaries even when the container has a read-only root filesystem (readOnlyRootFilesystem: true). Supporting noexec, nodev, and nosuid gives users a native way to harden volume mounts to match security benchmarks and policy.

The gap is most visible with emptyDir volumes, which are the most common writable volume type and have been the subject of multiple security findings:

  • Issue #48912: Recognized security gap - the inability to set mount options on emptyDir was flagged in an audit but remained unresolved until now.
  • Issue #119627: Kubernetes 1.24 Security Audit (Finding NCC-E003660-7HM) - external auditors specifically noted that the inability to mount emptyDir with noexec represents a security failure.

However, the same gap applies to all volume types. PersistentVolumes have a mountOptions field, but those options are filesystem-level flags applied by the CSI driver at the node, so they do not reliably translate into bind mount flags inside the container. Previously, there was no mechanism to set noexec, nosuid, or nodev on the bind mount that the container runtime creates for any volume type.

Additionally, the emptyDir volume type defaults to creating directories with a hardcoded mode of 0777. This previously meant that any process that can discover the volume could read, write, and delete anything in the volume, regardless of who created it.

You could - and still can - use an initial container to set a different access mode, but this is more complex, and hard to verify for compliance.

This causes real problems:

  • Multi-container pods sharing an emptyDir could not prevent one container from deleting another's files. The sticky bit (01777) solves this, but there was no native way to set it.
  • Some applications and security frameworks expect /tmp directories to have the sticky bit set (mode 01777). Without native support for setting the emptyDir mode, users had to use init containers or alternative volume types to meet this requirement.
  • Platform engineers who want tighter permissions (e.g., 0750 for owner and group only) have to use init containers running chmod, which adds unnecessary complexity.

The emptyDir volume type was a notable gap. As one of the most common writable volume types in Kubernetes, it had no way to control its creation permissions.

Real-world use cases

Application developers, working closely with security engineers, are responsible for maintaining the security posture of their applications and ensuring workloads do not pose risks to the wider infrastructure. These features allow development teams to confidently address critical security scenarios:

Preventing Privilege Escalation on Writable Mounts: An application developer configuring temporary workspace volumes (like emptyDir or /tmp mounts) can ensure they are mounted with nosuid and noexec. This guarantees that even if the application is compromised and a malicious payload is downloaded, the workload cannot execute the payload or use it to escalate privileges on the node.

Securing Shared Scratch Space in Multi-Container Pods: A developer configuring CI/CD pipeline pods often needs multiple containers (e.g., a builder container and a sidecar logger) to share a workspace. By setting mode: 01777 on an emptyDir, the developer ensures the shared workspace behaves like a traditional Unix /tmp directory. Each container can write files independently, but a compromised process in one container cannot delete the build artifacts produced by another.

Enforcing Principle of Least Privilege for Application Data: An application developer deploying a database pod can lock down access to the database's temporary storage. By setting mode: 0750 on the emptyDir, the developer ensures that only the specific database user and group can read or write to the volume, explicitly denying access to any other processes or sidecars in the same pod.

Note: Both features are behind Alpha feature gates in Kubernetes v1.37. To use them, enable VolumeBindMountOptions and EmptyDirVolumeMode on the API server and kubelet.

Example 1: Enforcing bind mount options

This full Pod manifest mounts an emptyDir volume at /tmp with bindMountOptions: [noexec, nosuid].

apiVersion: v1
kind: Pod
metadata:
  name: hardened-bindmount-pod
  namespace: default
spec:
  os:
    name: linux
  containers:
    - name: hardened-app
      image: alpine:latest
      command: ["sleep", "3600"]
      securityContext:
        readOnlyRootFilesystem: true
      volumeMounts:
        - name: temp-storage
          mountPath: /tmp
          bindMountOptions:
            - noexec
            - nosuid
  volumes:
    - name: temp-storage
      emptyDir: {}

Example 2: emptyDir volume permission mode with sticky bit

This full Pod manifest creates an emptyDir volume using mode: 01777 to enforce standard Unix /tmp sticky bit protections across containers.

apiVersion: v1
kind: Pod
metadata:
  name: hardened-emptydir-pod
  namespace: default
spec:
  os:
    name: linux
  containers:
    - name: app-container
      image: alpine:latest
      command: ["sleep", "3600"]
      volumeMounts:
        - name: shared-tmp
          mountPath: /tmp
  volumes:
    - name: shared-tmp
      emptyDir:
        mode: 01777

Verifying the features in Linux

To verify that these features are actively enforcing restrictions, you can run kubectl exec into the container. The following examples simulate attempts to perform actions that are successfully blocked by these features.

Verifying noexec

Attempt to write and run a script on a volume mounted with noexec:

# 1. Exec into the pod
kubectl exec -it hardened-bindmount-pod -- sh

# 2. Create an executable script on the mounted volume
cd /tmp
echo '#!/bin/sh' > test.sh
echo 'echo "Executing untrusted code..."' >> test.sh
chmod +x test.sh

# 3. Attempt to run the script
./test.sh

Expected result:

sh: ./test.sh: Permission denied

Even if an executable file is created, the Linux kernel refuses execution because MS_NOEXEC is enforced at the bind mount level.

Verifying the sticky bit

Attempt to delete another user's file in an emptyDir with 01777 permission mode:

# 1. Exec into the pod
kubectl exec -it hardened-emptydir-pod -- sh

# 2. Verify directory permissions on /tmp
ls -ld /tmp
# Output: drwxrwxrwt 2 root root ... /tmp (Notice the 't' indicating sticky bit)

# 3. Create a file as the guest user
su -s /bin/sh -c "touch /tmp/guest_file" guest

# 4. Attempt to delete that file as nobody
su -s /bin/sh -c "rm /tmp/guest_file" nobody

Expected result:

rm: can't remove '/tmp/guest_file': Operation not permitted

The kernel blocks deletion because the sticky bit (01777) restricts file removal strictly to the owner of the file.

Things to know

Keep these key details in mind as you begin using these features. Full details are available in the official documentation for bind mount options, emptyDir volume mode, and emptyDir volumes.

  • Default unchanged: If you omit bindMountOptions or do not set an emptyDir mode, you get standard default behaviors (like 0777 permissions) exactly as before.
  • Broad volume support: bindMountOptions works with emptyDir, PersistentVolumes, CSI volumes, projected volumes, ConfigMaps, Secrets, and more. The only exception is image volumes, which are explicitly unsupported. The mode field works with all emptyDir medium types: default (disk-backed), Memory (tmpfs), and HugePages.
  • Runtime capabilities matter (for bindMountOptions): The container runtime must support the CRI mount_options field and advertise it via runtimeFeatures. The scheduler uses node declared features to avoid placing pods on incompatible nodes. If a pod reaches such a node anyway, the kubelet rejects it. There is no silent degradation. However, using mode for an emptyDir does not require runtime support.
  • Not the same as PV mountOptions: PersistentVolume mountOptions apply at the storage layer via the CSI driver. The new bindMountOptions controls bind mount flags applied inside the container by the runtime. They operate at different layers and do not conflict.
  • fsGroup interaction: If fsGroup is set in the pod's security context, the group permissions applied by fsGroup will override the mode specified for the emptyDir volume. This is the same behavior that exists for defaultMode on Secret and ConfigMap volumes.
  • Linux only: Flags like noexec, nosuid, nodev, and Unix permission modes are Linux concepts. bindMountOptions has no effect on Windows nodes. On Windows, the mode field is also skipped for emptyDir volumes, since Windows does not support Unix-style file permissions.
  • Version skew safety: Both features are additive. For emptyDir mode: if the API server has the gate enabled but the kubelet does not, the field is accepted but ignored - the kubelet falls back to 0777. For bindMountOptions: the scheduler uses Node Declared Features to prevent placing pods on nodes without runtime support; if a pod reaches such a node, the kubelet rejects it rather than silently ignoring the options.
  • Feature Gates: Both capabilities are available as Alpha features in Kubernetes v1.37:
    • VolumeBindMountOptions: Controls bind mount flags on volume mounts.
    • EmptyDirVolumeMode: Controls creation permission modes on emptyDir volumes.

How do I get involved?

These new features are driven by SIG Node and SIG Storage. You can find more details in the KEPs for these enhancements: KEP-5855 (bind mount options) and KEP-5502 (emptyDir permission mode).

Reach out to SIG Node:

Reach out to SIG Storage:

Last modified September 11, 2026 at 9:58 AM PST: publish blogs for week 4 (42bb7682fe)