GPFS报错 “stale file handle”

最新推荐文章于 2025-02-27 20:36:44 发布

原创最新推荐文章于 2025-02-27 20:36:44 发布 · 1.6k 阅读

0 ·

CC 4.0 BY-SA版权

文章标签：

#linux #运维 #服务器

HPC 专栏收录该内容

7 篇文章

订阅专栏

本文介绍了如何应对计算机节点上GPFS服务因内存不足被杀死导致的文件系统无法使用的问题，包括检查、强制卸载并重新启动GPFS的详细步骤。解决方案简单易行，适合快速恢复服务。

部署运行你感兴趣的模型镜像

Unfortunately, the GPFS service running at compute nodes “mmfsd” gets killed sometimes by the out of memory killer. This usually unmounts GPFS from the compute node and makes it unavailable in LSF. The mount point /gpfs3 shows “stale file handle”. This can also happen even after the node was rebooted.

The good news is that there is a quite easy fix.

Check /gpfs3 mount point
[root@b35n20 ~]# ls -l /gpfs3/
ls: /gpfs3/: Stale file handle
ls: cannot open directory /gpfs3/: Stale file handle

shut down GPFS locally, start it again and wait a few seconds
[root@b35n20 ~]# mmshutdown
Wed Oct 2 09:32:10 CEST 2019: mmshutdown: Starting force unmount of GPFS file systems
Wed Oct 2 09:32:15 CEST 2019: mmshutdown: Shutting down GPFS daemons
Shutting down!
Unloading modules from /lib/modules/3.10.0-957.21.3.el7.x86_64/extra
Unloading module mmfs26
Unloading module mmfslinux
Wed Oct 2 09:32:19 CEST 2019: mmshutdown: Finished
[root@b35n20 ~]# mmstartup
Wed Oct 2 09:32:31 CEST 2019: mmstartup: Starting GPFS …

Check /gpfs3 mount point again
[root@b35n20 xcatpost]# ls -l /gpfs3
total 123534208
drwxrwxr-x 171 root root 32768 Sep 30 10:04 applications
…

您可能感兴趣的与本文相关的镜像

Stable-Diffusion-3.5

图片生成

Stable-Diffusion

Stable Diffusion 3.5 (SD 3.5) 是由 Stability AI 推出的新一代文本到图像生成模型，相比 3.0 版本，它提升了图像质量、运行速度和硬件效率