Parallel file system for AI + HPC
BeeGFS High-Performance Parallel Storage
Scale throughput, metadata, and capacity independently.
BeeGFS is a software-defined parallel file system engineered for highly concurrent AI, HPC, research, and data-intensive workloads. It stripes file data across multiple storage targets while distributing metadata across dedicated services to deliver aggregate performance from the entire cluster.
Applied Data Systems turns BeeGFS software into a complete storage platform—servers, NVMe and SSD media, network fabric, rack integration, tuning, validation, and support.
Request a Consultation

Workload fit
Where BeeGFS Fits Best
BeeGFS is strongest where many compute clients need concurrent access to a shared namespace and storage performance must scale with the compute cluster.
AI
Training & Checkpointing
Feed distributed GPU clusters with high-throughput training data and handle frequent checkpoint I/O without forcing all jobs through a single storage controller.
HPC
Simulation & Modeling
Parallel access for CFD, weather, physics, engineering, seismic, and MPI-oriented workflows that generate large concurrent I/O streams.
SCIENCE
Life Sciences & Research
Support genomics, microscopy, imaging, instrumentation, and research computing environments with shared scalable file access.
EDA
Electronic Design Automation
Scale metadata and file throughput independently for workflows that combine very large directory structures with high client concurrency.
ANALYTICS
Data-Intensive Analytics
Provide a fast shared POSIX data tier for analytics and scientific pipelines that need predictable access to large working datasets.
SCRATCH
High-Performance Scratch
Use BeeGFS as a dedicated scratch tier—or BeeOND for temporary job-local scratch—when transient I/O should not burden primary storage.
Design a BeeGFS Platform Around Your Workload
Bring ADS your compute architecture, file profile, throughput target, usable-capacity requirement, preferred network, and growth curve. We’ll turn those requirements into an integrated BeeGFS storage design.
Request a Consultation

Applied Data Systems
©2026 Applied Data Systems
How it works
Separate Metadata From the Data Path
BeeGFS separates metadata coordination from file-content storage. Clients obtain file placement information from metadata services, then communicate directly with multiple storage services for parallel I/O.
This architecture lets metadata performance and storage throughput scale independently while avoiding a central controller in the user-data path.
Add the resources the workload needs—metadata, storage capacity, or bandwidth—without redesigning the application environment.
Clients
Native Linux client provides the shared POSIX mount point.
Metadata
Distributed services coordinate namespace, permissions, and file placement.
Storage
File chunks are striped across multiple targets for parallel access.
Management
Tracks cluster configuration, nodes, targets, pools, mirrors, and state.
Why BeeGFS
Parallel Storage That Scales With the Cluster
BeeGFS is built around independent, distributed services. That lets organizations add performance, metadata capability, or capacity where the workload actually needs it instead of scaling every component at the same rate.
Scale-Out
Aggregate Performance
Files are striped across multiple storage targets so clients can access data from many servers in parallel and benefit from the bandwidth of the full storage cluster.
POSIX
Application Transparency
Applications see a familiar Linux file system and global namespace. Existing AI, HPC, simulation, and research applications do not need to be rewritten to use the parallel back end.
Open Choice
Hardware-Agnostic Design
Deploy BeeGFS on the server, flash, capacity media, and network technologies that fit the workload instead of being tied to a proprietary storage appliance.
Platform capabilities
Built for Modern AI and HPC Data Paths
BeeGFS combines a simple Linux file interface with technologies designed for high-throughput, low-overhead clustered I/O.
NVIDIA GPUDirect Storage
BeeGFS supports GPUDirect Storage to enable efficient data movement between storage, RDMA networking, and NVIDIA GPUs while reducing CPU involvement in the I/O path.
GDS
High-Speed Networking
Native RDMA support enables BeeGFS to take advantage of InfiniBand and RoCE fabrics for lower latency and reduced CPU overhead, with TCP/IP available as needed.
RDMA
Storage Pools
Group storage targets by performance or capacity characteristics and place workloads on the tier that best matches their I/O profile.
POOL
Mirroring & Resilience
Buddy mirroring capabilities can protect metadata and storage targets while keeping the parallel design distributed across the cluster.
HA
Data Movement
Remote Storage Targets and parallel copy capabilities support staging and movement between BeeGFS, POSIX file systems, and S3-compatible storage.
MOVE
Monitoring & Operations
System-wide monitoring, quota controls, event monitoring, and administrative tooling help operators manage large shared environments.
OPS
BeeGFS on Demand
Turn Local NVMe Into Per-Job Parallel Scratch
BeeOND creates temporary BeeGFS instances using SSD or NVMe devices inside compute nodes allocated to a job. The result is a high-performance local scratch tier that can absorb temporary and random I/O while reducing pressure on the global file system.
BeeOND integrates with cluster schedulers such as Slurm and can work alongside BeeGFS or other global storage platforms. Data can be staged into the job-local tier, processed at local-NVMe speed, then written back when the job completes.
Global Storage
Datasets + checkpoints
→
BeeOND
Job-local NVMe scratch
→
Compute
AI + HPC jobs
The ADS advantage
BeeGFS, Engineered as a Complete Storage System
BeeGFS gives you architectural freedom. ADS turns that freedom into a validated design by balancing storage media, CPU, PCIe lanes, network bandwidth, metadata services, rack density, and capacity around your workloads.
✓
Workload-Driven Architecture
Size metadata, data targets, media, and network based on file size, concurrency, throughput, capacity, and growth requirements.
✓
Best-of-Breed Hardware
Choose server platforms, NVMe, SSD, HDD, and fabric technologies that fit the technical and commercial requirements.
✓
Fabric Integration
Design around InfiniBand, RoCE, or Ethernet and validate the end-to-end data path between clients, metadata, and storage nodes.
✓
Factory Integration & Burn-In
Rack, cable, configure, tune, and test the full platform before it arrives at the customer site.
✓
GPU-Aware Storage Design
Integrate BeeGFS with GPU clusters and GPUDirect Storage-capable network paths where the application stack benefits from it.
✓
One Point of Accountability
ADS coordinates the delivered stack so customers do not have to troubleshoot storage, networking, and server vendors independently.
Design Priority
BeeGFS Approach
ADS Engineering Focus
Maximum Throughput
Stripe across more storage targets and servers.
NVMe density, PCIe topology, RDMA fabric, client-to-target balance.
Metadata Intensive
Scale metadata services independently.
Metadata node design, flash latency, namespace and directory behavior.
Capacity Growth
Add storage servers and targets without changing the application namespace.
Media mix, expansion domains, rack/power planning, lifecycle strategy.
AI / GPU I/O
Use parallel data access, RDMA, and GPUDirect Storage where appropriate.
GPU-to-NIC topology, fabric bandwidth, storage concurrency, validation.
