Development Guide¶
Contributing to and extending Mock NVML.
Prerequisites¶
- Go 1.25+ with CGo enabled
- GCC toolchain
- Docker (for cross-platform builds)
- golangci-lint (for linting)
Project Structure¶
pkg/gpu/mocknvml/
├── bridge/ # CGo bridge layer (hand-written + generated)
│ ├── cgo_types.go # Shared CGo type definitions
│ ├── helpers.go # Helper functions + main() + go:generate
│ ├── init.go # nvmlInit_v2, nvmlShutdown, etc.
│ ├── device.go # Device handle functions
│ ├── events.go # Event set/wait functions
│ ├── system.go # System functions
│ ├── internal.go # Internal export table (nvidia-smi)
│ └── stubs_generated.go # Auto-generated stubs (~289 functions)
├── engine/
│ ├── config.go # Configuration loading
│ ├── config_types.go # YAML struct definitions
│ ├── device.go # ConfigurableDevice implementation
│ ├── engine.go # Singleton engine
│ ├── handles.go # C↔Go handle mapping
│ ├── invalid_device.go # Invalid device handle sentinel
│ ├── utils.go # Debug utilities
│ ├── version.go # NVML version responses
│ └── *_test.go # Unit tests
├── configs/
│ ├── mock-nvml-config-a100.yaml
│ ├── mock-nvml-config-b200.yaml
│ ├── mock-nvml-config-gb200.yaml
│ ├── mock-nvml-config-gb300.yaml
│ ├── mock-nvml-config-h100.yaml
│ ├── mock-nvml-config-l40s.yaml
│ └── mock-nvml-config-t4.yaml
├── Dockerfile
├── Makefile
└── README.md
cmd/generate-bridge/
├── main.go # Stub generator (--stats, --validate flags)
├── parser.go # nvml.h prototype parser
└── main_test.go # Generator tests
tests/mocknvml/
├── bridge_tests.go # Bridge-level integration tests
├── main.go # Integration test (mini device plugin)
├── Dockerfile
├── Makefile
└── README.md
docs/
├── README.md
├── quickstart.md
├── architecture.md
├── configuration.md
├── cuda-mock.md
├── development.md
├── examples.md
├── troubleshooting.md
├── demo/
│ ├── standalone/
│ └── with-fgo/
└── integrations/
└── fake-gpu-operator.md
Building¶
Local Build¶
Docker Build (Cross-Platform)¶
Build with Custom Version¶
Clean Build¶
Running Tests¶
Unit Tests¶
With Coverage¶
Integration Test¶
Adding New NVML Functions¶
Step 1: Identify the Function¶
Find the function signature in go-nvml:
// In github.com/NVIDIA/go-nvml/pkg/nvml
type Device interface {
GetNewFunction() (ReturnType, Return)
}
Step 2: Implement in ConfigurableDevice¶
Add the method to pkg/gpu/mocknvml/engine/device.go:
// GetNewFunction returns the new function value
func (d *ConfigurableDevice) GetNewFunction() (ReturnType, nvml.Return) {
// Check if config provides this value
if d.config != nil && d.config.NewProperty != nil {
value := d.config.NewProperty.Value
debugLog("[NVML] nvmlDeviceGetNewFunction -> %v\n", value)
return value, nvml.SUCCESS
}
// No config = not supported
debugLog("[NVML] nvmlDeviceGetNewFunction -> NOT_SUPPORTED\n")
return ReturnType{}, nvml.ERROR_NOT_SUPPORTED
}
Step 3: Add Config Types (if needed)¶
Add to pkg/gpu/mocknvml/engine/config_types.go:
// NewPropertyConfig defines the new property configuration
type NewPropertyConfig struct {
Value int `json:"value,omitempty"`
Enabled bool `json:"enabled,omitempty"`
}
Add field to DeviceConfig:
type DeviceConfig struct {
// ... existing fields ...
NewProperty *NewPropertyConfig `json:"new_property,omitempty"`
}
Step 4: Add Bridge Implementation¶
Add the C-exported function to the appropriate bridge file (e.g., bridge/device.go
for device functions, or create a new file for a new category):
//export nvmlDeviceGetNewFunction
func nvmlDeviceGetNewFunction(nvmlDevice unsafe.Pointer, result unsafe.Pointer) C.nvmlReturn_t {
if result == nil {
return C.NVML_ERROR_INVALID_ARGUMENT
}
dev := engine.GetEngine().LookupDevice(uintptr(nvmlDevice))
value, ret := dev.GetNewFunction()
if ret == nvml.SUCCESS {
*(*C.int)(result) = C.int(value)
}
return toReturn(ret)
}
Step 5: Regenerate Stubs¶
The generator automatically detects your new implementation and removes its stub:
Or from repo root:
Step 6: Test¶
Adding New GPU Profiles¶
Step 1: Create YAML File¶
Create pkg/gpu/mocknvml/configs/mock-nvml-config-newgpu.yaml:
version: "1.0"
system:
driver_version: "560.35.03"
nvml_version: "12.560.35.03"
cuda_version: "12.6"
cuda_version_major: 12
cuda_version_minor: 6
device_defaults:
name: "NVIDIA New GPU"
architecture: "hopper" # or appropriate arch
# ... add all properties
devices:
- index: 0
uuid: "GPU-ae000000-0000-0000-0000-000000000000"
pci:
bus_id: "0000:07:00.0"
# ... add all devices
Step 2: Test the Profile¶
Step 3: Verify All nvidia-smi Commands¶
# Basic display
nvidia-smi
# Full query
nvidia-smi -q
# XML output
nvidia-smi -x -q
# Specific queries
nvidia-smi -q -d MEMORY
nvidia-smi -q -d POWER
nvidia-smi -q -d TEMPERATURE
Code Style¶
Go Style¶
Follow standard Go conventions:
Documentation¶
- Public functions must have doc comments
- Comments should be ≤80 characters per line
- Use
debugLog()for debug output, notfmt.Printf
Testing¶
- Every new function should have unit tests
- Test both success and error cases
- Use table-driven tests where appropriate
func TestNewFunction(t *testing.T) {
tests := []struct {
name string
config *DeviceConfig
expected ReturnType
wantErr nvml.Return
}{
{
name: "with config",
config: &DeviceConfig{NewProperty: &NewPropertyConfig{Value: 42}},
expected: 42,
wantErr: nvml.SUCCESS,
},
{
name: "without config",
config: nil,
expected: 0,
wantErr: nvml.ERROR_NOT_SUPPORTED,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
dev := &ConfigurableDevice{config: tt.config}
got, ret := dev.GetNewFunction()
if ret != tt.wantErr {
t.Errorf("GetNewFunction() ret = %v, want %v", ret, tt.wantErr)
}
if got != tt.expected {
t.Errorf("GetNewFunction() = %v, want %v", got, tt.expected)
}
})
}
}
Debugging¶
Enable Debug Logging¶
Debug with GDB¶
# Build with debug symbols
CGO_CFLAGS="-g" go build -gcflags="all=-N -l" -buildmode=c-shared ...
# Attach GDB
gdb nvidia-smi
(gdb) set environment LD_LIBRARY_PATH .
(gdb) run
Debug with Delve¶
Regenerating Stubs¶
The stub generator creates stubs_generated.go with stub implementations for
NVML functions that don't have hand-written implementations:
# Preferred: use Makefile target from repo root
make generate
# Or from bridge directory (uses go:generate directive)
cd pkg/gpu/mocknvml/bridge
go generate
# Or run generator directly with all flags
go run ./cmd/generate-bridge \
-input vendor/github.com/NVIDIA/go-nvml/pkg/nvml/nvml.go \
-header vendor/github.com/NVIDIA/go-nvml/pkg/nvml/nvml.h \
-bridge pkg/gpu/mocknvml/bridge \
-output pkg/gpu/mocknvml/bridge/stubs_generated.go
The generator:
1. Parses all NVML function names from nvml.go
2. Parses C prototypes from nvml.h for ABI-correct signatures
3. Scans bridge/*.go for existing //export directives
4. Generates stubs only for functions NOT already implemented
5. Outputs to stubs_generated.go
When you add a new implementation to a bridge file, regenerate stubs to automatically remove the corresponding stub.
Coverage Statistics¶
View current implementation coverage:
Output:
NVML Function Coverage:
Total functions: 396
Hand-written implementations: 107 (27.0%)
Generated stubs: 289 (73.0%)
By file:
device.go: 94 functions
events.go: 6 functions
init.go: 5 functions
system.go: 4 functions
helpers.go: 1 function
internal.go: 1 function
Signature Validation¶
Check that hand-written exports match nvml.h parameter counts:
This compares the number of Go parameters in each //export function against
the corresponding C prototype in nvml.h. Exits non-zero on mismatch — useful
in CI to catch signature drift.
Release Checklist¶
- [ ] All tests pass:
go test -race ./... - [ ] Linter passes:
golangci-lint run - [ ] Integration test passes:
make -C tests/mocknvml test - [ ] Documentation updated
- [ ] YAML configs updated if needed
- [ ] README updated with new features
- [ ] Version bumped if applicable
Contributing¶
- Fork the repository
- Create a feature branch
- Make changes with tests
- Run linter and tests
- Submit pull request
Commit Message Format¶
Types: feat, fix, docs, test, refactor, chore