Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ jobs:
with:
python-version: ${{ matrix.python }}
cache: pip
- run: python -m pip install -e '.[dev,viz]'
- run: python -m pip install -e '.[dev]'
- if: matrix.backend == 'numba-cuda'
# The device backend imports CUDA runtime metadata even for CPU-only tests.
# Only the runtime is needed here; simulation does not require a GPU/driver.
Expand Down
8 changes: 5 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -1,16 +1,18 @@
# Contributing

Install `.[dev,viz]`; run `pytest`, `ruff check .`, and `ruff format --check .`.
Install `.[dev]`; run `pytest`, `ruff check .`, and `ruff format --check .`.
See [tests/README.md](tests/README.md) for CPU, simulator, and hardware commands.

Keep boundaries clear:

- `src/gpu_backtest/core/` contains generic computation only. It must not import
examples, test references, benchmark code, CLI, workflows, or deployment tools.
- `workflows/` uses core primitives for splitting/analysis/output charts.
examples, test references, benchmark code, CLI, or deployment tools.
- Example trading rules/data stay under `examples/` and helper code under `tools/`.
- CLI modules are thin adapters; they do not own numerical or cloud logic.

Keep output raw: no built-in scoring, ranking, statistical inference, splitting or charts.
Preserve four grouped array values and their parameter alignment.

New example strategies need a device-function spec and hand-calculated case.
Use synthetic data. Keep credentials, private plugins/datasets, and research output
out of Git. New tracked files require a reviewed update to the explicit allowlist
Expand Down
115 changes: 66 additions & 49 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,12 @@

## 10× faster on our public billion-pair benchmark

**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** Same RSI grid:
**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** The same public RSI grid:
**1,000,000,000 pairs × 1,024 bars**, compared with an **eight-thread compiled
Numba CPU baseline**. Measured speedup **10.34×**, saving about seven minutes per
sweep. Cloud setup is additional; [method and raw evidence](docs/benchmarks.md).
Numba CPU baseline**. Measured speedup **10.34×** on v0.4, saving about seven
minutes per sweep. Cloud setup is additional; [method and raw evidence](docs/benchmarks.md).
v0.6 preserves the GPU reductions and saves raw results instead of ranked analysis.
The historical timing includes v0.4's output processing; it is not a new v0.6 timing.

- **Your algorithm:** load a strategy module/object; private rules can remain private.
- **Large grids:** deterministic GPU reductions without materializing the full return matrix.
Expand All @@ -15,28 +17,26 @@ sweep. Cloud setup is additional; [method and raw evidence](docs/benchmarks.md).

```text
src/gpu_backtest/
core/ GPU engine, kernels, grids, indicators, statistics, output
workflows/ Generic split/common analysis and charts
core/ GPU engine, kernels, grids, indicators, raw output
cli/ Thin command-line adapters
examples/gpu_backtest_examples/
rsi/ Example strategy + config + generated CSV
data.py Example/benchmark synthetic data generator
rsi/ Educational strategy + config + generated CSV
data.py Synthetic data generator
tools/gpu_backtest_tools/
runpod/ Optional API / SSH / bundle / lifecycle helper
benchmarks/ Performance runner and compiled CPU baseline
checks/ Hardware smoke checks and CPU test reference
benchmarks/ Performance runner and compiled CPU comparison
checks/ Hardware smoke checks and CPU numeric reference
tests/
cpu/ CPU contracts, examples, workflows, helper tests
cpu/ Contracts, examples, CLI, helpers and CPU reference
gpu/ Isolated CUDA simulation and real-GPU tests
docs/ Strategy contract, RunPod usage, benchmark method
benchmarks/results/ Historical measurements and validation evidence
```

**Start with `core/engine.py`** for the GPU run. Numerical kernels are in
**Start with `core/engine.py`** for a GPU run. CUDA kernels are in
`core/kernels.py`; trading rules are supplied by a plugin. Core imports no example,
CPU comparison engine, benchmark, or RunPod code. CPU preprocessing of market data
and indicator tables is part of the GPU pipeline; the separate CPU backtest baseline
is only a benchmark/reference tool.
CPU backtest comparison, benchmark, or RunPod code. Market validation and indicator
preparation use the CPU before GPU execution.

## Try the RSI example without a GPU

Expand All @@ -45,17 +45,13 @@ Python 3.11–3.13, from a checkout:
```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev,viz]'
python -m pip install -e '.[dev]'
NUMBA_ENABLE_CUDASIM=1 gpu-backtest run \
--config examples/gpu_backtest_examples/rsi/config.json --out-prefix runs/rsi
gpu-backtest charts --input runs/rsi_top_entry.csv runs/rsi_top_exit.csv \
--output runs/rsi.html
```

The simulator is for tiny examples/tests. Open `runs/rsi.html` in a browser;
its Vega libraries load from a public CDN. The [RSI example](examples/gpu_backtest_examples/rsi/README.md)
is educational and uses generated data. It is packaged separately as
`gpu_backtest_examples.rsi.strategy`, not as an engine builtin.
The CUDA simulator is for tiny examples/tests. The [RSI example](examples/gpu_backtest_examples/rsi/README.md)
uses generated data and is packaged separately as `gpu_backtest_examples.rsi.strategy`.

## Use your own strategy

Expand All @@ -69,43 +65,64 @@ results = run(example, "market.csv", "runs/custom", buy=0.0015, sell=0.0015)
```

Or use `gpu-backtest run --strategy my_strategies.example --input market.csv
--out-prefix runs/custom`. Plugins implement a complete Numba CUDA trading loop;
see [the contract](docs/strategy.md). Entry/exit parameters are separate Cartesian
axes (up to four dimensions each), with up to four precomputed tables. Shared-parameter
diagonal-only sweeps are not supported. The API returns grouped sums/squares and
optional ranked artifacts, not a trade ledger, equity curve, Sharpe, or drawdown series.
--out-prefix runs/custom`. Plugins implement their complete Numba CUDA trading loop;
see [the contract](docs/strategy.md). Entry/exit parameters form separate Cartesian
axes, up to four dimensions each, with up to four precomputed indicator tables.
Shared-parameter diagonal-only sweeps are not supported.

## Optional workflows and helpers
## Raw output

| Command | Owner / purpose |
A run returns four float64 arrays and, when an output prefix is supplied, writes:

- `<prefix>_results.npz`: `entry_sum`, `entry_sumsq`, `exit_sum`, `exit_sumsq`.
- `<prefix>_manifest.json`: ordered parameter dimensions, counts, fees, array schema,
result-file SHA-256, and reduction consistency checks.

Entry arrays aggregate each entry parameter set across **all exits**; exit arrays
aggregate each exit set across **all entries**. Array indices follow Cartesian
parameter order, with the last dimension varying fastest. Individual returns and
their squares round to float32 before float64 accumulation. No full pair matrix,
trade ledger or equity curve is stored.

```python
import numpy as np

with np.load("runs/rsi_results.npz", allow_pickle=False) as results:
entry_sum = results["entry_sum"]
exit_sum = results["exit_sum"]
```

Apply your own analysis downstream. v0.6 removes built-in scoring, ranking,
common-parameter analysis, date-split pipelines and charts. The `top_n` option and
old ranked CSV/schema-1 artifacts are removed; numerical return arrays remain the
same. Old config keys are rejected rather than silently ignored.

[RTX 4090 validation](benchmarks/results/rtx4090_raw_output_parity_20261005.json):
v0.5/v0.6 returned arrays matched byte-for-byte on small, 16,777,216-pair and
1,000,000,000-pair grids. Clean-wheel results matched the default RunPod bundle,
all five physical-GPU tests passed, and owned-pod deletion was confirmed.

## RunPod and validation tools

| Command | Purpose |
|---|---|
| `pipeline` | `workflows/`: split, run each segment, intersect ranked parameters |
| `split`, `common`, `charts` | `workflows/`: standalone data/result analysis |
| `runpod` | `tools/runpod/`: lease one GPU, upload selected files, download, delete |
| `benchmark` | `tools/benchmarks/`: reproduce the public CPU/GPU measurements |
| `gpu-check` | `tools/checks/`: small numeric checks on actual hardware |
| `run` | One GPU backtest with raw results |
| `runpod` | Lease one GPU, upload selected files, run, download, delete |
| `benchmark` | CPU/GPU performance and numeric comparison |
| `gpu-check` | Small numerical checks on actual hardware |

```bash
gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \
--ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts
--ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi
```

RunPod requires your account/API key and registered SSH key. It is an optional
execution helper; [setup, manual GPU route, and cleanup](docs/runpod.md).
Commands keep their existing names. The public `from gpu_backtest import run` API
and old `rsi_meanrev` shorthand remain usable; direct internal imports moved under
`core/`, `workflows/`, or `gpu_backtest_tools` in v0.5.

## Performance and development

| Same billion-pair RSI job | CPU, eight threads | RTX 4090 | Saved per sweep |
|---|---|---|---|
| 1,000,000,000 pairs × 1,024 bars, through output | 7 min 44.60 s | 44.95 s | 6 min 59.65 s; 10.34× faster |
RunPod requires your account/API key and registered SSH key. It defaults to one
full-input run; [setup, manual GPU route, and cleanup](docs/runpod.md).
The public `from gpu_backtest import run` API and old `rsi_meanrev` shorthand
remain usable. Direct internal imports moved under `core/` or
`gpu_backtest_tools` in v0.5.

This is one measured run per side on the published setup, not a universal speed
claim. Small CPU jobs may not justify cloud startup. Statistics describe parameter
combination groups, not independent market samples or a forecast of profits.
See [benchmark details](docs/benchmarks.md) for scope and raw data.
## Development

```bash
python -m pytest
Expand Down
Loading
Loading