RAID6 on a PERC H730P: 6×4TB backup server layout and tuning
Why we chose RAID6 for a 6-disk backup server on a Dell PERC H730P, how we built the virtual disk in F2 HII, LVM + XFS for /data, and the tuning and SMART checks.
- The server would temporarily hold the only copy of a NAS’s photo archive, so we wanted any two disks to be able to fail. That meant RAID6, not the RAID10 our production server uses.
- We planned for 8 disks. Two turned out bad (one was dead, the other had an amber fault LED that followed it to another slot), so the array is RAID6 on 6×4 TB: 14.55 TiB, with
/dataon XFS at 13.6 TiB. - After install, SMART showed two disks at about 74,300 power-on hours, one of them with 2 grown defects. Two-disk redundancy was the right choice.
- A post-install script handles the tuning: writeback limits, read-ahead, BBR, CPU governor, Docker log caps, journald cap, fail2ban, and smartd entries for each disk behind the controller (
-d megaraid,N).
We set up a Dell PowerEdge R730 as a home backup server. Its first job was to receive a full copy of a NAS (about 5.8 TB) so the NAS could be reformatted. During that window, this server would hold the only copy of years of photos. This post covers the disk decisions, the controller settings, the filesystem layout, and the tuning script, using the numbers we actually measured. The usernames and paths in the code are placeholders.
Why RAID6 instead of RAID10
Our production server is a comparable Dell machine with the same PERC H730P controller and the same 4 TB SAS drive model, and it runs RAID10. For a database host that’s a good choice. For this box, we compared the two:
| RAID10 (production server) | RAID6 (backup server) | |
|---|---|---|
| Failures survived | 1 guaranteed; 2 only if they hit different mirror pairs | any 2 |
| Usable space, 6×4 TB | ~3 disks’ worth | 4 disks’ worth |
| Sequential read/write | good | good (striped, controller write-back cache with battery) |
| Small random writes | very good | weaker, which matters little for a backup target |
The deciding factor was rebuild risk. Rebuilding a 4 TB disk takes hours. If a second disk fails during the rebuild and it happens to be the mirror partner of the first, RAID10 loses data. RAID6 survives any second failure, and during the period when this box held the only copy, that’s what mattered. The extra usable space was a bonus.
We made that decision before we could see the disks’ history. Once Debian was running, smartctl showed it was justified:
for i in 0 1 2 3 4 5; do
printf "disk%s " $i
sudo smartctl -H -A -d megaraid,$i /dev/sda | grep -E \
"SMART Health|grown defect|Current Drive Temperature|power on time" | tr "\n" "|"
echo
done
| Disk | SMART health | Power-on hours | Grown defect list | Temp |
|---|---|---|---|---|
| 0 | OK | 11,870 | 0 | 40 °C |
| 1 | OK | 14,312 | 0 | 40 °C |
| 2 | OK | 14,312 | 0 | 40 °C |
| 3 | OK | 11,869 | 0 | 41 °C |
| 4 | OK | 74,314 | 2 | 39 °C |
| 5 | OK | 74,280 | 0 | 40 °C |
Two of the six disks have been powered on for about eight and a half years. One of them has already reallocated two sectors. That doesn’t mean it will fail tomorrow, but this is the array where you want two-disk tolerance. A stable grown-defect count is fine. A count that keeps rising means the disk needs replacing.
Finding the bad disks before building the array
The server had eight drive bays, and the plan was RAID6 across all eight. Physical Disk Management in the controller showed only seven disks. Six were Ready. The bay holding the missing disk had no LEDs lit, and a seventh disk showed as Foreign (left over from a previous array) in a bay with an amber LED. We narrowed it down by moving disks between bays. That’s safe to do before any virtual disk exists.
- The disk that hadn’t been detected at all showed up as Failed, 0 KB once moved to the other bay. The disk itself was dead, and the bays were fine.
- The Foreign disk was recognized normally (3.638 TB), but its amber LED moved with it to the other bay. On Dell servers, an amber drive LED usually indicates a fault or predicted failure, so we didn’t trust it.
We removed both and built the array from the six disks that showed Ready. Pull the suspect disks rather than just leaving them unticked. That way Check All can’t add one by accident. Usable capacity dropped from the planned ~21.8 TiB to 14.55 TiB, which was still more than double the data we had to hold.
Building the virtual disk in F2 System Setup (HII)
Instead of the older Ctrl+R utility, we used the HII path in System Setup, which is what Dell’s PERC 9 user guide documents:
F2 → Device Settings → Integrated RAID Controller 1: Dell PERC H730P Mini Configuration Utility → Configuration Management → Create Virtual Disk
| Setting | Value | Why |
|---|---|---|
| RAID level | RAID6 | two-disk tolerance |
| Physical disks | Check All (the 6 Ready disks) → Apply Changes | |
| Size | the maximum it fills in | one virtual disk, LVM splits it |
| Strip element size | 256 KB (default was 64 KB) | large sequential files: video, camera RAW |
| Read policy | Read Ahead | sequential backup reads |
| Write policy | Write Back | already preselected; we confirmed afterwards that the controller battery reports good |
| Disk cache | Disable | the drives’ own caches are not battery-backed |
| Default initialization | Fast | the background initialization finishes later |
Clear the old metadata from a Foreign disk first, under Configuration Management → Manage Foreign Configuration → Clear Foreign Configuration. After creating the disk, Virtual Disk Management should show RAID6 and Optimal. From Linux, megasasctl later reported PERC H730P Mini … batt:good and 14TiB RAID 6 1x6 optimal, with all six member disks online.
LVM and XFS for /data
The OS was installed by a preseeded Debian 13 ISO, which is covered in the unattended install post. It put everything on LVM on the RAID virtual disk:
| Mount | Size (as installed) | Filesystem |
|---|---|---|
/boot/efi |
976 MiB | vfat |
/boot |
977 MiB | ext4 |
/ |
95.4 GiB | ext4, noatime |
/var |
476.8 GiB | ext4, noatime (Docker images and logs) |
| swap | 30.5 GiB | |
/tmp |
47.7 GiB | ext4, noatime |
/home |
286.1 GiB | ext4, noatime |
/data |
13.6 TiB | XFS, noatime |
We chose XFS for /data because it was going to hold well over a hundred thousand files, from multi-gigabyte videos down to small JPEGs, and be written by parallel copy jobs. On top of the installer’s noatime, the post-install script sets logbsize=256k and inode64 in fstab. The xfs(5) man page lists 256k as the largest log buffer size for version 2 logs, and inode64 has been the default since kernel 3.7, so listing it is just explicit. One thing we hit in testing: a remount did not change logbsize. A full unmount and mount did. After a reboot, the server showed rw,noatime,attr2,inode64,logbufs=8,logbsize=256k,noquota.
The thing we got wrong: the partitioning recipe capped /data at 18 TB so the planned eight-disk array would keep a few TiB of free space in the volume group for LVM snapshots, as a restore point before bulk photo reorganization. With six disks, 18 TB was more than the space left, so /data took everything and vgs shows VFree 0. XFS can grow but can’t shrink, so getting that headroom back means rebuilding the volume or adding disks.
The tuning script
The post-install script is idempotent, so running it again gives the same result. It started from our production server’s settings and adds a few for bulk storage. Main pieces:
# /etc/sysctl.d/90-backup.conf (excerpt)
vm.swappiness = 10
vm.dirty_background_bytes = 268435456 # start background writeback at 256 MiB
vm.dirty_bytes = 2147483648 # make writers flush at 2 GiB
vm.vfs_cache_pressure = 50 # keep dentry/inode caches longer
vm.min_free_kbytes = 1048576
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr
net.core.netdev_max_backlog = 16384
fs.file-max = 2097152
The machine has 128 GB of RAM. With percentage-based writeback limits, a large copy can collect gigabytes of dirty pages and then stall while they flush. Byte limits keep each flush small. The kernel documentation notes that dirty_bytes and dirty_ratio are alternatives: setting one disables the other. A lower vfs_cache_pressure makes the kernel slower to drop cached directory and inode entries, which helps when tools walk a large photo tree again and again.
# /etc/udev/rules.d/60-backup-block.rules: the PERC virtual disk reports vendor DELL
ACTION=="add|change", KERNEL=="sd[a-z]", ATTRS{vendor}=="DELL*", \
ATTR{queue/scheduler}="mq-deadline", ATTR{queue/read_ahead_kb}="4096", ATTR{queue/nr_requests}="1024"
# and the same 4 MiB read-ahead on the LVM volumes (8192 sectors)
for d in /dev/mapper/backup--vg-*; do blockdev --setra 8192 "$d"; done
Other pieces of the script:
- CPU governor
performance, set by a small oneshot systemd unit (ExecStart=-/usr/bin/cpupower frequency-set -g performance; the leading-keeps the unit from failing on hardware without frequency scaling). nofileraised to 1,048,576 in/etc/security/limits.d/.- Docker log caps.
/etc/docker/daemon.jsonsets"log-driver": "json-file"with"max-size": "50m", "max-file": "3", pluslive-restore. Docker’s defaultmax-sizeis unlimited. Our production server has nodaemon.json, so its container logs can grow without bound. Changing the daemon setting only affects containers created afterwards. - journald capped with
SystemMaxUse=2G. - fail2ban and vnstat enabled.
On the real server, the script exited with status 0. We checked the effects directly: the disk queue showed none [mq-deadline] with ra=4096, sysctl reported bbr, 2147483648 and 10, Docker reported overlay2 and cgroup v2, and docker, fail2ban, smartmontools and vnstat were all active.
SMART monitoring behind the MegaRAID controller
Linux sees the RAID virtual disk as /dev/sda, and the physical disks sit behind the controller. The smartctl(8) man page documents how to reach them: -d megaraid,N addresses physical disk N behind a MegaRAID-family controller, and Dell PERC is one of those. Instead of relying on a bare DEVICESCAN line, our script writes one explicit smartd line per disk that smartctl --scan reports, so each disk gets its own schedule:
if megasasctl 2>/dev/null | grep -q 'PERC'; then
{ echo "# physical disks behind the PERC"
for i in $(smartctl --scan | awk '/megaraid,/{sub(/.*megaraid,/,"",$3); print $3}'); do
echo "/dev/sda -d megaraid,$i -a -s (S/../.././02|L/../../6/03)"
done; } > /etc/smartd.conf
else
echo "DEVICESCAN -a -s (S/../.././02|L/../../6/03)" > /etc/smartd.conf
fi
systemctl restart smartmontools
The -s regular expression is matched against T/MM/DD/d/HH, as described in smartd.conf(5). Here that means a short self-test every day in the 02:00 hour and a long test on Saturdays (6) in the 03:00 hour. An earlier version listed eight disks by hand, which would have been wrong for the six-disk array. Deriving the list from smartctl --scan fixes that, but only when the script runs, so run it again after adding or replacing disks.
On the server, megasasctl reported the PERC, so the script took the MegaRAID branch, and smartmontools came up active. In the test VM, which has no SMART-capable disks, it had stayed inactive. We didn’t print the generated /etc/smartd.conf during setup, so check yours:
cat /etc/smartd.conf # expect one /dev/sda -d megaraid,N line per disk
sudo smartctl --scan # what the scan found
journalctl -u smartmontools -b # smartd logs each device it registered
There’s one limitation to know about: these lines have no -m directive, so smartd logs problems but sends no notifications. Our production server sends a daily RAID and SMART summary from a separate script. Adding a similar check here, one that also alerts when the grown-defect count increases, is still on our list.
What we’d do differently
- Base the
/datamaximum on the array you actually get, or leave a fixed amount of free space in the volume group. A cap based on the plan cost us snapshot space. - Read SMART before choosing the RAID level if you can. We chose correctly, but by reasoning about rebuild risk, not from the data.
- Pull suspect disks out of the chassis before using Check All.
- Still to do: a hot spare in one of the empty bays, a replacement for the disk with grown defects, and confirming the controller’s patrol read schedule. We discussed all three and haven’t done any of them yet.
The server is on its own network segment and can reach only the NAS. That setup is in the RouterOS DMZ post. The copy itself is in migrating a Synology NAS to a Linux server with rsync.