eMMC storage set to read-only

I’ve recently had some very serious trouble with the Ceph cluster on my cluster box. While trying to debug it, I found out that the internal eMMC storage of two of the nodes has been set to read-only for whatever reason. I get such messages in the syslog:

[145619.057561] systemd-journald[412]: Failed to rotate /var/log/journal/ef1ec9435b634c2283e071eb0bcb4d4c/user-1000.journal: Read-only file system
[145619.062304] systemd-journald[412]: Failed to write entry to /var/log/journal/ef1ec9435b634c2283e071eb0bcb4d4c/system.journal (22 items, 739 bytes) despite vacuuming, ignoring: Read-only file system (Dropped 1 similar message(s))

The well-known tools don’t say that the FS is read-only:

mixtile@blade3n3:~$ df -h
Filesystem      Size  Used Avail Use% Mounted on
tmpfs           1.6G  2.4M  1.6G   1% /run
/dev/mmcblk0p3  115G   12G   99G  11% /
tmpfs           7.8G     0  7.8G   0% /dev/shm
tmpfs           5.0M   20K  5.0M   1% /run/lock
tmpfs           1.6G  120K  1.6G   1% /run/user/1000
mixtile@blade3n3:~$ cat /etc/fstab
# <file system>     <mount point>  <type>  <options>   <dump>  <fsck>
UUID=a5e5be9e-44a0-4d14-9ea4-83aed9de6634 /              ext4    defaults,x-systemd.growfs    0       1
mixtile@blade3n3:~$ sudo mount -o remount,rw /
[sudo] password for mixtile: 
mount: /: cannot remount /dev/mmcblk0p3 read-write, is write-protected.
       dmesg(1) may have more information after failed mount system call.
mixtile@blade3n3:~$ nano TESTDATEI
Unable to create directory /home/mixtile/.local/share/nano/: Read-only file system
It is required for saving/loading search history or cursor positions.
mixtile@blade3n3:~$ lsblk
NAME         MAJ:MIN RM   SIZE RO TYPE MOUNTPOINTS
loop0          7:0    0    69M  1 loop /snap/core22/2956
loop1          7:1    0  54.2M  1 loop /snap/core26/463
loop2          7:2    0 106.2M  1 loop /snap/lxd/40423
loop3          7:3    0    69M  1 loop /snap/core22/2438
loop4          7:4    0 106.3M  1 loop /snap/lxd/40918
loop5          7:5    0  38.6M  1 loop /snap/snapd/28255
loop6          7:6    0  43.5M  1 loop /snap/snapd/27740
mmcblk0      179:0    0 116.5G  0 disk 
├─mmcblk0p1  179:1    0     4M  0 part 
├─mmcblk0p2  179:2    0   512B  0 part 
└─mmcblk0p3  179:3    0 116.5G  0 part /
mmcblk0boot0 179:32   0     4M  1 disk 
mmcblk0boot1 179:64   0     4M  1 disk 
nvme0n1      259:0    0   7.5T  0 disk 

Rebooting the nodes does not help BTW. The device tree appears to be fresh:

-rwxr-xr-x 1 root root 270451 Sep 22 17:33 rk3588-mixtile-blade3.dtb

Could you share the output of the following commands from one of the affected Blade 3 nodes?

# Check MMC, ext4 and I/O errors
sudo dmesg -T | grep -Ei 'mmc|ext4|I/O error|read-only|write.protect|timeout|reset'

# Check actual root filesystem mount options
findmnt -no SOURCE,FSTYPE,OPTIONS /

# Check block device read-only status
cat /sys/block/mmcblk0/ro

# Check eMMC health and lifetime
sudo mmc extcsd read /dev/mmcblk0 | grep -Ei 'LIFE_TIME|PRE_EOL|WRITE_PROTECT'

I’m particularly interested in whether this is an ext4 filesystem issue, eMMC wear, or an RK3588 MMC controller/driver problem.

Since two nodes are affected, it would also be interesting to know whether they started failing at approximately the same time.

mixtile@blade3n4:~$ sudo dmesg -T | grep -Ei 'mmc|ext4|I/O error|read-only|write.protect|timeout|reset'
[sudo] password for mixtile: 
[Thu Oct  8 16:58:30 2026] rk806 spi2.0: no reset-setting pinctrl state
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: No normal pinctrl state
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: No idle pinctrl state
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: IDMAC supports 32-bit address mode.
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: Using internal DMA controller.
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: Version ID is 270a
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: DW MMC controller at irq 91,32 bit host data width,256 deep fifo
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: Looking up vmmc-supply from device tree
[Thu Oct  8 16:58:30 2026] sdhci-dwcmshc fe2e0000.mmc: Looking up vmmc-supply from device tree
[Thu Oct  8 16:58:30 2026] sdhci-dwcmshc fe2e0000.mmc: Looking up vmmc-supply property in node /mmc@fe2e0000 failed
[Thu Oct  8 16:58:30 2026] sdhci-dwcmshc fe2e0000.mmc: Looking up vqmmc-supply from device tree
[Thu Oct  8 16:58:30 2026] sdhci-dwcmshc fe2e0000.mmc: Looking up vqmmc-supply property in node /mmc@fe2e0000 failed
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: Looking up vqmmc-supply from device tree
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: Failed getting OCR mask: -22
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: could not set regulator OCR (-22)
[Thu Oct  8 16:58:30 2026] dwmmc_rockchip fe2c0000.mmc: failed to enable vmmc regulator
[Thu Oct  8 16:58:30 2026] mmc_host mmc1: Bus speed (slot 0) = 400000Hz (slot req 400000Hz, actual 400000HZ div = 0)
[Thu Oct  8 16:58:31 2026] mmc0: SDHCI controller on fe2e0000.mmc [fe2e0000.mmc] using ADMA
[Thu Oct  8 16:58:31 2026] mpp_vdpu2 fdb50400.vdpu: reset_group->rw_sem_on=0
[Thu Oct  8 16:58:31 2026] mpp_vdpu2 fdb50400.vdpu: reset_group->rw_sem_on=0
[Thu Oct  8 16:58:31 2026] rkvdec2_init:1197: No niu aclk reset resource define
[Thu Oct  8 16:58:31 2026] rkvdec2_init:1200: No niu hclk reset resource define
[Thu Oct  8 16:58:31 2026] rkvdec2_init:1197: No niu aclk reset resource define
[Thu Oct  8 16:58:31 2026] rkvdec2_init:1200: No niu hclk reset resource define
[Thu Oct  8 16:58:31 2026] mali fb000000.gpu: Capping CSF_FIRMWARE_TIMEOUT to CSF_FIRMWARE_PING_TIMEOUT
[Thu Oct  8 16:58:31 2026] mmc0: Host Software Queue enabled
[Thu Oct  8 16:58:31 2026] mmc0: new HS200 MMC card at address 0001
[Thu Oct  8 16:58:31 2026] mmcblk0: mmc0:0001 DA4128 116 GiB 
[Thu Oct  8 16:58:31 2026]  mmcblk0: p1 p2 p3
[Thu Oct  8 16:58:31 2026] mmcblk0boot0: mmc0:0001 DA4128 4.00 MiB 
[Thu Oct  8 16:58:31 2026] mmcblk0boot1: mmc0:0001 DA4128 4.00 MiB 
[Thu Oct  8 16:58:31 2026] mmcblk0rpmb: mmc0:0001 DA4128 16.0 MiB, chardev (235:0)
[Thu Oct  8 16:58:36 2026] EXT4-fs (mmcblk0p3): mounted filesystem with ordered data mode. Quota mode: none.
[Thu Oct  8 16:58:37 2026] EXT4-fs (mmcblk0p3): re-mounted. Quota mode: none.
[Thu Oct  8 16:58:37 2026] EXT4-fs (mmcblk0p3): resizing filesystem from 30530552 to 30530552 blocks
mixtile@blade3n4:~$ findmnt -no SOURCE,FSTYPE,OPTIONS /
/dev/mmcblk0p3 ext4   rw,relatime
mixtile@blade3n4:~$ cat /sys/block/mmcblk0/ro
0
mixtile@blade3n4:~$ sudo mmc extcsd read /dev/mmcblk0 | grep -Ei 'LIFE_TIME|PRE_EOL|WRITE_PROTECT'
eMMC Life Time Estimation A [EXT_CSD_DEVICE_LIFE_TIME_EST_TYP_A]: 0x01
eMMC Life Time Estimation B [EXT_CSD_DEVICE_LIFE_TIME_EST_TYP_B]: 0x01
eMMC Pre EOL information [EXT_CSD_PRE_EOL_INFO]: 0x01

Supposedly yes, but after another reboot of the whole cluster, the write protection has disappeared. But: This issue has already led to serious trouble with my Ceph cluster and the data on it.

Thanks for the logs.

The eMMC health indicators look normal (0x01), with no signs of significant wear.

I noticed some MMC regulator initialization errors, but they belong to fe2c0000.mmc, while the eMMC is connected to fe2e0000.mmc.

Since two nodes experienced the issue around the same time and a full cluster reboot restored write access, I suspect a possible platform-level issue involving power management, the MMC controller, or the kernel rather than eMMC wear.

Could you also check the ext4 error counters and previous boot logs?

sudo tune2fs -l /dev/mmcblk0p3 | grep -Ei 'Filesystem state|Errors behavior|FS Error count|First error|Last error'
sudo journalctl -k -b -1 --no-pager | grep -Ei 'mmc|sdhci|ext4|I/O error|timeout|read-only'
uname -a

It would be useful to identify the root cause, especially since this has already affected Ceph data integrity.

One thing puzzles me: Why is eMMC speed still set to 400 MHz although I reduced it to 200 MHz in the device tree?

Filesystem state:         clean
Errors behavior:          Continue
Oct 06 22:00:00 blade3n3 kernel: sdhci: Secure Digital Host Controller Interface driver
Oct 06 22:00:00 blade3n3 kernel: sdhci: Copyright(c) Pierre Ossman
Oct 06 22:00:00 blade3n3 kernel: sdhci-pltfm: SDHCI platform and OF driver helper
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: No normal pinctrl state
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: No idle pinctrl state
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: IDMAC supports 32-bit address mode.
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: Using internal DMA controller.
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: Version ID is 270a
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: DW MMC controller at irq 91,32 bit host data width,256 deep fifo
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: Looking up vmmc-supply from device tree
Oct 06 22:00:00 blade3n3 kernel: sdhci-dwcmshc fe2e0000.mmc: Looking up vmmc-supply from device tree
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: Looking up vqmmc-supply from device tree
Oct 06 22:00:00 blade3n3 kernel: sdhci-dwcmshc fe2e0000.mmc: Looking up vmmc-supply property in node /mmc@fe2e0000 failed
Oct 06 22:00:00 blade3n3 kernel: sdhci-dwcmshc fe2e0000.mmc: Looking up vqmmc-supply from device tree
Oct 06 22:00:00 blade3n3 kernel: sdhci-dwcmshc fe2e0000.mmc: Looking up vqmmc-supply property in node /mmc@fe2e0000 failed
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: Failed getting OCR mask: -22
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: could not set regulator OCR (-22)
Oct 06 22:00:00 blade3n3 kernel: dwmmc_rockchip fe2c0000.mmc: failed to enable vmmc regulator
Oct 06 22:00:00 blade3n3 kernel: mmc_host mmc1: Bus speed (slot 0) = 400000Hz (slot req 400000Hz, actual 400000HZ div = 0)
Oct 06 22:00:00 blade3n3 kernel: mmc0: SDHCI controller on fe2e0000.mmc [fe2e0000.mmc] using ADMA
Oct 06 22:00:00 blade3n3 kernel: mali fb000000.gpu: Capping CSF_FIRMWARE_TIMEOUT to CSF_FIRMWARE_PING_TIMEOUT
Oct 06 22:00:00 blade3n3 kernel: mmc0: Host Software Queue enabled
Oct 06 22:00:00 blade3n3 kernel: mmc0: new HS200 MMC card at address 0001
Oct 06 22:00:00 blade3n3 kernel: mmcblk0: mmc0:0001 DA4128 116 GiB 
Oct 06 22:00:00 blade3n3 kernel:  mmcblk0: p1 p2 p3
Oct 06 22:00:00 blade3n3 kernel: mmcblk0boot0: mmc0:0001 DA4128 4.00 MiB 
Oct 06 22:00:00 blade3n3 kernel: mmcblk0boot1: mmc0:0001 DA4128 4.00 MiB 
Oct 06 22:00:00 blade3n3 kernel: mmcblk0rpmb: mmc0:0001 DA4128 16.0 MiB, chardev (235:0)
Oct 06 22:00:00 blade3n3 kernel: EXT4-fs (mmcblk0p3): mounted filesystem with ordered data mode. Quota mode: none.
Oct 06 22:00:00 blade3n3 kernel: EXT4-fs (mmcblk0p3): re-mounted. Quota mode: none.
Oct 06 22:00:00 blade3n3 kernel: EXT4-fs (mmcblk0p3): resizing filesystem from 30530552 to 30530552 blocks
Linux blade3n3 6.1.0-1027-rockchip #27 SMP Sun Apr 27 01:54:34 UTC 2025 aarch64 aarch64 aarch64 GNU/Linux

The log actually shows 400 kHz, not 400 MHz. That’s the initial MMC initialization clock.

Also, this message comes from fe2c0000.mmc (mmc1), while your eMMC is connected to fe2e0000.mmc (mmc0).

Your eMMC is detected in HS200 mode, which supports up to 200 MHz.

Could you check the actual eMMC clock?

sudo cat /sys/kernel/debug/mmc0/ios

This should show both the requested and actual clock frequencies.

Also, the ext4 filesystem reports a clean state, and the previous boot logs don’t show any obvious MMC I/O errors. So we still don’t have evidence explaining why the filesystem became read-only.

clock:          200000000 Hz
actual clock:   200000000 Hz
vdd:            18 (3.0 ~ 3.1 V)
bus mode:       2 (push-pull)
chip select:    0 (don't care)
power mode:     2 (on)
bus width:      3 (8 bits)
timing spec:    9 (mmc HS200)
signal voltage: 1 (1.80 V)
driver type:    0 (driver type B)

So at last, it’s 200 MHz, not 400, right? I was really shocked when reading the previous syslog message.

Yes, exactly! Your eMMC is running at 200 MHz in HS200 mode with an 8-bit bus.

The earlier 400000 Hz message was just the 400 kHz initialization clock of a different MMC controller (mmc1), not your eMMC (mmc0).

So your device tree frequency setting appears to be working correctly.