Parallel file system for AI + HPC

BeeGFS High-Performance Parallel Storage

Scale throughput, metadata, and capacity independently.

BeeGFS is a software-defined parallel file system engineered for highly concurrent AI, HPC, research, and data-intensive workloads. It stripes file data across multiple storage targets while distributing metadata across dedicated services to deliver aggregate performance from the entire cluster.

Applied Data Systems turns BeeGFS software into a complete storage platform—servers, NVMe and SSD media, network fabric, rack integration, tuning, validation, and support.

Request a Consultation

Workload fit

Where BeeGFS Fits Best

BeeGFS is strongest where many compute clients need concurrent access to a shared namespace and storage performance must scale with the compute cluster.

AI

Training & Checkpointing

Feed distributed GPU clusters with high-throughput training data and handle frequent checkpoint I/O without forcing all jobs through a single storage controller.

HPC

Simulation & Modeling

Parallel access for CFD, weather, physics, engineering, seismic, and MPI-oriented workflows that generate large concurrent I/O streams.

SCIENCE

Life Sciences & Research

Support genomics, microscopy, imaging, instrumentation, and research computing environments with shared scalable file access.

EDA

Electronic Design Automation

Scale metadata and file throughput independently for workflows that combine very large directory structures with high client concurrency.

ANALYTICS

Data-Intensive Analytics

Provide a fast shared POSIX data tier for analytics and scientific pipelines that need predictable access to large working datasets.

SCRATCH

High-Performance Scratch

Use BeeGFS as a dedicated scratch tier—or BeeOND for temporary job-local scratch—when transient I/O should not burden primary storage.

Design a BeeGFS Platform Around Your Workload

Bring ADS your compute architecture, file profile, throughput target, usable-capacity requirement, preferred network, and growth curve. We’ll turn those requirements into an integrated BeeGFS storage design.

Request a Consultation

Applied Data Systems

©2026 Applied Data Systems

How it works

Separate Metadata From the Data Path

BeeGFS separates metadata coordination from file-content storage. Clients obtain file placement information from metadata services, then communicate directly with multiple storage services for parallel I/O.

This architecture lets metadata performance and storage throughput scale independently while avoiding a central controller in the user-data path.

Add the resources the workload needs—metadata, storage capacity, or bandwidth—without redesigning the application environment.

Clients

Native Linux client provides the shared POSIX mount point.

Metadata

Distributed services coordinate namespace, permissions, and file placement.

Storage

File chunks are striped across multiple targets for parallel access.

Management

Tracks cluster configuration, nodes, targets, pools, mirrors, and state.

Why BeeGFS

Parallel Storage That Scales With the Cluster

BeeGFS is built around independent, distributed services. That lets organizations add performance, metadata capability, or capacity where the workload actually needs it instead of scaling every component at the same rate.

Scale-Out

Aggregate Performance

Files are striped across multiple storage targets so clients can access data from many servers in parallel and benefit from the bandwidth of the full storage cluster.

POSIX

Application Transparency

Applications see a familiar Linux file system and global namespace. Existing AI, HPC, simulation, and research applications do not need to be rewritten to use the parallel back end.

Open Choice

Hardware-Agnostic Design

Deploy BeeGFS on the server, flash, capacity media, and network technologies that fit the workload instead of being tied to a proprietary storage appliance.

Platform capabilities

Built for Modern AI and HPC Data Paths

BeeGFS combines a simple Linux file interface with technologies designed for high-throughput, low-overhead clustered I/O.

NVIDIA GPUDirect Storage

BeeGFS supports GPUDirect Storage to enable efficient data movement between storage, RDMA networking, and NVIDIA GPUs while reducing CPU involvement in the I/O path.

GDS

High-Speed Networking

Native RDMA support enables BeeGFS to take advantage of InfiniBand and RoCE fabrics for lower latency and reduced CPU overhead, with TCP/IP available as needed.

RDMA

Storage Pools

Group storage targets by performance or capacity characteristics and place workloads on the tier that best matches their I/O profile.

POOL

Mirroring & Resilience

Buddy mirroring capabilities can protect metadata and storage targets while keeping the parallel design distributed across the cluster.

HA

Data Movement

Remote Storage Targets and parallel copy capabilities support staging and movement between BeeGFS, POSIX file systems, and S3-compatible storage.

MOVE

Monitoring & Operations

System-wide monitoring, quota controls, event monitoring, and administrative tooling help operators manage large shared environments.

OPS

BeeGFS on Demand

Turn Local NVMe Into Per-Job Parallel Scratch

BeeOND creates temporary BeeGFS instances using SSD or NVMe devices inside compute nodes allocated to a job. The result is a high-performance local scratch tier that can absorb temporary and random I/O while reducing pressure on the global file system.

BeeOND integrates with cluster schedulers such as Slurm and can work alongside BeeGFS or other global storage platforms. Data can be staged into the job-local tier, processed at local-NVMe speed, then written back when the job completes.

Global Storage

Datasets + checkpoints

BeeOND

Job-local NVMe scratch

Compute

AI + HPC jobs

The ADS advantage

BeeGFS, Engineered as a Complete Storage System

BeeGFS gives you architectural freedom. ADS turns that freedom into a validated design by balancing storage media, CPU, PCIe lanes, network bandwidth, metadata services, rack density, and capacity around your workloads.

Workload-Driven Architecture

Size metadata, data targets, media, and network based on file size, concurrency, throughput, capacity, and growth requirements.

Best-of-Breed Hardware

Choose server platforms, NVMe, SSD, HDD, and fabric technologies that fit the technical and commercial requirements.

Fabric Integration

Design around InfiniBand, RoCE, or Ethernet and validate the end-to-end data path between clients, metadata, and storage nodes.

Factory Integration & Burn-In

Rack, cable, configure, tune, and test the full platform before it arrives at the customer site.

GPU-Aware Storage Design

Integrate BeeGFS with GPU clusters and GPUDirect Storage-capable network paths where the application stack benefits from it.

One Point of Accountability

ADS coordinates the delivered stack so customers do not have to troubleshoot storage, networking, and server vendors independently.

Design Priority

BeeGFS Approach

ADS Engineering Focus

Maximum Throughput

Stripe across more storage targets and servers.

NVMe density, PCIe topology, RDMA fabric, client-to-target balance.

Metadata Intensive

Scale metadata services independently.

Metadata node design, flash latency, namespace and directory behavior.

Capacity Growth

Add storage servers and targets without changing the application namespace.

Media mix, expansion domains, rack/power planning, lifecycle strategy.

AI / GPU I/O

Use parallel data access, RDMA, and GPUDirect Storage where appropriate.

GPU-to-NIC topology, fabric bandwidth, storage concurrency, validation.