Crash Recovery & Inode Persistence¶
This document details the crash recovery and state persistence mechanisms of DVFS, covering the theoretical durability model (In Essence) and the engineering solutions for restart survival (Implementation Quirks).
1. In Essence: Durability & High Availability¶
A distributed filesystem must survive unexpected node crashes, power outages, and service restarts without data corruption, ghost entries, or desynchronized client sessions.
DVFS achieves durability and fault isolation through three core invariants:
1. Authoritative FileServer State Durability: Physical data, directory hierarchies, permissions, quotas, and inode identifiers persist on physical disk. A restarted FileServer reconstructs its exact pre-crash state without relying on the MetaServer.
2. Deterministic Identity Persistence: File Identifiers (FIDs) must be stable across reboots. If a client holds a reference to fs1_42_1, the FileServer must resolve fs1_42_1 to the exact same file after a restart.
3. Heartbeat Monitoring & Stale Isolation: The MetaServer continuously monitors storage nodes via heartbeats. If a node goes offline, the MetaServer immediately isolates it from new client routing requests while preserving existing ownership mappings for when the node returns.
[Running Cluster] ---> [FileServer Node Crashes] ---> [MetaServer Ticker Exceeds 30s]
|
v
[Client Receives Error] <--- [Routing Excludes Node] <--- [Marked Stale]
|
v
[Node Reboots & Loads State] ---> [Registration & Heartbeat] ---> [Marked Healthy]
|
v
[Client Resumes Ops]
2. Implementation Quirks & Practical Realities¶
2.1 The Persistent InodeStore (.dvfs_inodes_index.json)¶
The Problem: Volatile Inode IDs¶
In early DVFS prototypes, FileServer startup walked the directory tree in parallel goroutines and assigned Inode IDs sequentially from 0. This created severe bugs:
- Depending on thread scheduling, on one boot a folder was assigned ID 2, and on the next boot it received ID 5.
- If a client was connected during a fileserver restart, running refresh sent ListDir using the old FID (fs1_2_1), which now pointed to an entirely different or empty folder.
- The client's cache handler wiped out its local cached directory tree because the server returned an empty directory listing.
The Solution: InodeStore Architecture¶
In internal/fileserver/inodestore.go, every FileServer root maintains a persistent mapping file named .dvfs_inodes_index.json:
- Path Normalization: Before lookup or storage, paths are converted to forward slashes (
/) and cleaned viafilepath.Cleanto guarantee cross-platform consistency between Linux and Windows hosts. - Deterministic ID Allocation: If a path exists in
.dvfs_inodes_index.json, its Inode ID is preserved across restarts. Only new files allocate a fresh ID fromnext_inode_id. - Thread Safety: All InodeStore operations are synchronized with a dedicated
sync.Mutex. - Atomic Commits: Saved via write-to-temp and atomic
os.Rename.
Client Session Restoration (ReRegister)¶
When a FileServer crashes and restarts, its volatile in-memory client registrations and callback channels are erased. Rather than forcing clients to terminate and relaunch:
1. When a client executes refresh (or on subsequent directory access), CacheHandler.Refresh() invokes Client.ReRegister().
2. The client re-sends RegisterClient RPC with its active ClientID, CallbackAddress, Username, RootUser, and RootPath.
3. The restarted FileServer re-establishes the callback session and returns its root FID.
4. The client validates that the root FID returned matches its cached root FID (ensured by .dvfs_inodes_index.json), seamlessly restoring push invalidations.
2.2 MetaServer State Recovery (MongoDB)¶
Updated: MetaServer state moved from
metaserver_state.jsonto MongoDB. The collections arefileservers(keyed on the node's stable-id),users, andshares(keyed on grantee+owner+path). Recovery is a singleLoadSnapshotread at boot; the in-memory maps remain the read path forGetRootsandNavigate. There is no import path from the old file — the cluster starts empty and rebuilds itself as fileservers register. The JSON shape below is retained for historical reference only.
The MetaServer coordinator used to persist its complete operational state in metaserver_state.json:
{
"fileservers": {
"0": {
"address": "10.7.52.85:50052",
"user_count": 2,
"last_heartbeat_unix": 1773062400,
"status": "healthy"
}
},
"users": {
"alice": 0,
"bob": 0
},
"shared": {
"bob": [
{
"Owner": "alice",
"Path": "alice/shared_proj",
"DisplayName": "shared_proj"
}
]
},
"next_fs_id": 1
}
- Startup Reconstitution:
NewMetaServer(*stateFile)used to parse the JSON snapshot at launch. It now takes aMetaStoreand issues a singleLoadSnapshotread, restoring known fileservers, user-to-fileserver assignments, and shared directory registries into the in-memory maps. A store that cannot be read is fatal rather than a warning: starting with empty routing state would strand every user. - Heartbeat Evaluation: Unchanged. Heartbeat timestamps are evaluated against the current
Unix time, and a node that has not checked in within the timeout window transitions to
stale. - ID Allocation:
next_fs_idwas an auto-incrementing integer in the JSON file. Numeric ids are now allocated from thecounterscollection, and a node's identity is its stable-id(fs1) rather than its address, so a DHCP lease change no longer registers the same machine twice. - Unregistered Home Nodes: A user whose home node is absent from the snapshot is retained
as orphaned rather than dropped. They get no live route, and
GetRootsrefuses to assign them a new home node, because reassignment would overwrite their stored placement and strand whatever data is still on the original node. The account becomes routable again as soon as that node registers. The boot log reports the count asorphaned_users. An operator who knows the node is gone for good releases its users through the admin console's remove-node API, which callsDeregisterFileServer. Because an orphan's home node has no record, it does not appear in the console's node list; release it by its stable id instead (DELETE /api/nodes/<fs_id>, e.g./api/nodes/fs3). Released users are placed afresh on their next login, and whatever data remained on the old node is no longer reachable through DVFS.
2.3 Heartbeat & Stale Transitions¶
- FileServer Heartbeat Loop: A background goroutine in
internal/fileserver/msclient.goexecutesHeartbeat(address)to the MetaServer every 5 seconds (configurable via-meta_heartbeat_interval). - MetaServer Monitor Ticker: A background goroutine runs every 5 seconds (
-heartbeat_check_interval): - Checks each registered server's
now - LastHeartbeatUnix. - If delta > 30 seconds (
-heartbeat_timeout), the node status transitions fromhealthytostale. - Routing Exclusion: When a client requests
Navigate(username, rootUser), the MetaServer checks the hosting node's status. Stale nodes are rejected with:"root user 'alice' is currently unavailable". - Automatic Re-Attachment: When an offline FileServer comes back online, its registration loop automatically reconnects to the MetaServer, sends a
Heartbeat, and the MetaServer restores its status tohealthy.
2.4 Admin Console Cold-Start Recovery¶
The Admin Console maintains historical monitoring graphs and alerts that survive restarts:
1. Ring-Buffer Snapshot (admin_metrics_snapshot.json):
- Every 60 seconds and during graceful shutdown (Stop()), SaveMetricsSnapshot flushes the 60-minute ring buffers (720 data points per node) to disk.
- On startup, LoadMetricsSnapshot reloads historical telemetry so graphs do not reset to zero.
2. Alert Engine Persistence (admin_alerts.json):
- Active and resolved alerts are persisted via atomic file writes, maintaining an audit trail across console restarts.
3. Crash Recovery Verification Runbooks¶
3.1 Four-Terminal MDS Crash Recovery Test¶
This test verifies that the MetaServer coordinator persists its routing table, restores state on cold reboot, and allows client operations to continue without FileServer restarts.
Terminal Layout¶
- Terminal 1 (MDS): MetaServer
- Terminal 2 (FS): FileServer
fs1 - Terminal 3 (Client A): Client
alice - Terminal 4 (Client B): Client
bob(verification)
Step 1: Start MetaServer¶
./bin/metaserver \
-port=50051 \
-mongo_uri=mongodb://127.0.0.1:27017/dvfs \
-heartbeat_timeout=30s \
-heartbeat_check_interval=5s
Step 2: Start FileServer¶
./bin/fileserver \
-id=fs1 \
-port=50052 \
-data=./fileserver_data/fs1 \
-meta_addr=127.0.0.1:50051 \
-own_ip=127.0.0.1 \
-meta_retry_interval=1s \
-meta_heartbeat_interval=2s
Step 3: Establish Client State & Snapshot¶
In Terminal 3, launch Client A:
Inside the client REPL, upload a test file: In Terminal 4 (optional), launch Client B to populate another user: Confirm the routing state is written to MongoDB:Step 4: Crash and Restart MetaServer¶
- In Terminal 1, stop the MetaServer with
Ctrl+C(orkill -9). - Keep the FileServer running in Terminal 2. Note in FS logs that retry/heartbeat attempts temporarily fail.
- Restart the MetaServer against the same database:
- Observe the logs:
- MDS Logs: State recovery log reporting restored counts (
fileservers,users,shares,orphaned_users). - FS Logs: Re-connection and heartbeat success logs resume within the retry interval.
Step 5: Validate Resumption¶
Start a new client session for alice:
test.txt without restarting the FileServer.
3.2 Heartbeat & Stale Transition Test (Fast-Test Flags)¶
To verify the MetaServer's liveness tracker and stale node isolation without waiting 30+ seconds, run with accelerated heartbeat flags:
Step 1: Start MetaServer with Fast Heartbeat Checks¶
./bin/metaserver \
-port=50051 \
-mongo_uri=mongodb://127.0.0.1:27017/dvfs \
-heartbeat_timeout=6s \
-heartbeat_check_interval=1s
Step 2: Start FileServer with High-Frequency Heartbeats¶
./bin/fileserver \
-id=fs1 \
-port=50052 \
-data=./fileserver_data/fs1 \
-meta_addr=127.0.0.1:50051 \
-own_ip=127.0.0.1 \
-meta_retry_interval=1s \
-meta_heartbeat_interval=2s
healthy.
Step 3: Simulate Crash & Observe Stale Transition¶
- Kill the FileServer process (
Ctrl+Corkill -9). - Wait 6 seconds (
-heartbeat_timeout=6s). - Observe MDS logs:
node fs1 heartbeat timed out; status transitioned to stale. - Try connecting a new client:
The client receives an immediate failure:
root user 'alice' is on unavailable file server.
Step 4: Restart FileServer & Verify Auto Re-Attachment¶
- Restart the FileServer with the same
-datapath. - FileServer loads
.dvfs_inodes_index.json, registers with MDS, and sends a Heartbeat. - MDS logs confirm status transitioned back to
healthy. - Existing client runs
refresh(triggeringReRegister()) or new clients connect seamlessly.
3.3 FileServer Restart & Client Session Re-Registration Test¶
This test verifies that persistent inode allocation (.dvfs_inodes_index.json) and client ReRegister() prevent cache wipeouts when a storage node restarts.
Step 1: Populate Client Session¶
- Launch Client as
aliceand upload test files:
Step 2: Restart Storage Node¶
Restart the FileServer process while leaving the client session active:
# Via systemd
sudo systemctl restart dvfs-fileserver
# Or kill and relaunch binary with the exact same -data directory
Step 3: Trigger Client Refresh¶
In the active dvfs> client prompt, execute:
Expected Verification Results¶
- No Stale FID Cache Wipeout:
lsdisplaysprojectsandreport.pdfimmediately without returning an empty directory. - Session Re-Registration: The client calls
ReRegister()behind the scenes, restoring push invalidation callbacks on the restarted FileServer. - Index Stability: Verify
.dvfs_inodes_index.jsonin the FileServer-dataroot retains the exact same inode IDs for existing paths without reallocation.
3.4 Automated Regression Tests¶
To run the automated test suite for MetaServer crash recovery and state persistence:
Diagrams¶
MetaServer Crash Recovery State Machine¶
stateDiagram-v2
state "Starting" as Starting
state "LoadingState" as LoadingState
state "Ready" as Ready
state "Serving" as Serving
state "Crashed" as Crashed
state "Recovering" as Recovering
[*] --> Starting
Starting --> LoadingState : "LoadSnapshot from MongoDB"
LoadingState --> Ready : "snapshot hydrated (may be empty)"
LoadingState --> [*] : "store unreachable (fatal: exit)"
Ready --> Serving : "Start gRPC server"
Serving --> Crashed : "process killed"
Crashed --> Recovering : "process restarted"
Recovering --> LoadingState : "reload snapshot"