GPFS CES Configuration and Backup⚓︎
This page collects practical notes for working with IBM Storage Scale Cluster Export Services (CES), with a focus on service checks, SMB configuration handling, and backup planning.
What CES Does⚓︎
When a CES node leaves the GPFS cluster, CES IP addresses assigned to that node can be redistributed to other healthy CES nodes. The exact behavior depends on the configured address distribution policy and current service state.
CES is commonly used for:
- NFS exports
- SMB shares
- protocol service failover
- protocol configuration management
Quick CES Checks⚓︎
These are useful first checks when validating CES behavior or investigating a failover issue.
Show CES cluster configuration⚓︎
Show CES service state⚓︎
Show CES addresses⚓︎
Show CES log level⚓︎
Show CES component state across the cluster⚓︎
List CES events⚓︎
Move a CES IP manually⚓︎
CES Troubleshooting Notes⚓︎
If CES health information appears inconsistent:
If mmhealth metadata appears stale or out of sync, a resync may help:
Use that carefully and confirm cluster health again afterward.
SMB Configuration in CES⚓︎
In CES deployments, SMB uses GPFS-managed service components such as:
gpfs-smbgpfs-winbind
Important Configuration Rule⚓︎
Make changes in the main Samba configuration source used for import, not in generated GPFS-managed runtime files that may be overwritten later.
For example:
- source config:
/etc/samba/smb.conf - generated or managed config:
/var/mmfs/ces/smb.conf
Do not edit the generated file directly unless you fully understand the regeneration workflow.
Example SMB Configuration Pattern⚓︎
The exact values will vary by environment, but a typical Active Directory backed CES SMB configuration includes:
- AD workgroup and realm settings
- idmap backend configuration
- Kerberos settings
- default home directory and shell templates
- encryption settings
- one or more managed share definitions
Applying SMB Configuration Changes⚓︎
After updating the source configuration:
Validate the config⚓︎
Import it into the CES-managed configuration⚓︎
Restart CES SMB services⚓︎
Restart the services in the correct order so name mapping and SMB state come back cleanly:
A compact one-liner looks like:
net conf import /etc/samba/smb.conf && systemctl restart gpfs-smb && systemctl restart gpfs-winbind
SMB Logs and Diagnostics⚓︎
Useful log files include:
/var/adm/ras/log.smbd
/var/adm/ras/log.smbd.old
/var/adm/ras/log.winbindd
/var/adm/ras/log.winbindd-idmap
/var/adm/ras/log.winbindd-dc-connect
Useful commands:
tail -F /var/adm/ras/log.smbd
tail -F /var/adm/ras/log.winbindd
tail -F /var/adm/ras/log.winbindd-idmap
tail -F /var/adm/ras/log.winbindd-dc-connect
grep smbd /var/log/messages
grep winbindd /var/log/messages
Cache location example⚓︎
Preparing a New SMB Share⚓︎
Linux-side preparation often looks like this:
Create the share directory⚓︎
Set owner and group⚓︎
chown <user> /gpfs/<filesystem>/smbgroup/<share-name>
chgrp <group> /gpfs/<filesystem>/smbgroup/<share-name>
Set permissions⚓︎
chmod o-rx /gpfs/<filesystem>/smbgroup/<share-name>
chmod g+ws /gpfs/<filesystem>/smbgroup/<share-name>
CES and Protocol Configuration Backup⚓︎
Backups for a GPFS environment should include more than file data alone. In practice, administrators usually protect at least these categories:
- Cluster configuration data
- File system configuration data
- File system contents
- Protocol configuration data
Why CCR matters⚓︎
When the Cluster Configuration Repository (CCR) is used, the master copy of configuration data is stored redundantly across quorum nodes instead of relying on a separate primary or backup configuration server.
That improves resilience for GPFS administrative metadata as long as a quorum of CCR participants remains available.
Cluster Configuration Backup⚓︎
At a minimum, save:
- the output of
mmlscluster - a CCR backup or
mmsdrfs-equivalent repository backup, depending on repository type - any operational snapshots or supporting restore data your site depends on
Record cluster configuration⚓︎
CCR Backup and Restore⚓︎
If your environment uses an mmsdrbackup user exit workflow, document and automate it clearly.
Backup⚓︎
Example:
The implementation details depend on how your site configured the backup wrapper around the user exit.
Restore⚓︎
Before using mmsdrrestore, confirm:
- the backup file is valid
- the target nodes are correct
- you understand the recovery scope
File System Backup with SOBAR⚓︎
For disaster recovery scenarios, GPFS file system configuration and image backup workflows often use:
mmbackupconfigmmimgbackupmmrestoreconfigmmimgrestore
Important ordering⚓︎
Run:
mmbackupconfig- snapshot creation
mmimgbackup
For restore:
mmrestoreconfig- mount read-only if required
mmimgrestore- restore quotas if needed
- remount read-write
What mmbackupconfig captures⚓︎
The backup file can include items such as:
- NSD and disk configuration
- storage pools
- filesets and junctions
- policy rules
- quota definitions and limits
- file system attributes
It does not back up normal user file data.
Example SOBAR Backup Workflow⚓︎
Back up file system configuration⚓︎
Create a global snapshot⚓︎
Back up the image⚓︎
Example SOBAR Restore Workflow⚓︎
Optional: generate a report file for recreation planning⚓︎
Restore essential configuration⚓︎
Mount read-only if required for image restore⚓︎
Restore the image⚓︎
Unmount after restore⚓︎
Restore quotas if needed⚓︎
Mount read-write again⚓︎
Recommended Backup Checklist⚓︎
- export
mmlsclusteroutput regularly - maintain a tested CCR backup workflow
- back up SMB/NFS protocol configuration in a repeatable way
- document where imported Samba config is stored
- include
mmbackupconfigoutputs in file system backup procedures - pair image backups with consistent snapshots
- periodically test restore steps in a non-production environment
References⚓︎
- IBM Storage Scale Administration Guide sections on CES, CCR, SMB, and SOBAR
- IBM Spectrum Scale and Linux compatibility matrix