How We Reduced CI Time by 33% with Runtime-Based Test Splitting
A practical guide to optimizing parallel test execution in GitHub Actions
33%
Time Reduction
21%
Cost Savings
83%
DB Setup Faster
The Problem
Our CI pipeline was taking over 10 minutes to complete, with RSpec tests running across 7 parallel nodes. Despite the parallelization, we noticed significant imbalances between nodes - some finishing in 6 minutes while others took over 10 minutes.
The root causes were:
- 1. Filesize-based splitting - Tests were distributed by file size, not execution time
- 2. Redundant DB setup - Each node was creating 7 databases when only 1 was needed
- 3. Over-provisioned nodes - More nodes than necessary for the workload
Before vs After Comparison
Slowest: 10m 12s · Variance: 4m+
Slowest: 7m 24s · Variance: 1m
The Solution
1. Runtime-Based Splitting
Instead of distributing tests by file size, we switched to runtime-based splitting using parallel_tests (a gem for running tests in parallel across multiple CPU cores) and its runtime log feature.
How it works:
- 1. Each node records test execution times to a runtime log
- 2. Logs are uploaded as artifacts after each run
- 3. A merge job combines logs and caches the result
- 4. Next run uses cached logs for optimal distribution
- name: Run tests
run: |
RUNTIME_LOG="tmp/parallel_runtime_rspec.log"
if [ -f "$RUNTIME_LOG" ]; then
GROUP_BY="runtime"
else
GROUP_BY="filesize"
fi
bundle exec parallel_test spec -t rspec \
--group-by "$GROUP_BY" \
--runtime-log "$RUNTIME_LOG" 2. DB Setup Optimization
We discovered that each CI node was creating 7 databases during setup, even though each node has its own isolated MySQL service and only uses 1 database. Our project uses ridgepole (a DSL-based schema management tool) instead of Rails migrations.
Before (110s)
parallel:create[7] + ridgepole:parallel_prepare After (19s)
db:create + ridgepole:apply 3. Node Count Reduction
With better test distribution and faster DB setup, we reduced from 7 to 5 nodes while maintaining acceptable wall time. This yielded a 21% reduction in compute costs.
Results
| Metric | Before | After | Change |
|---|---|---|---|
| Wall Time | 10m 12s | 6m 50s | -33% |
| DB Setup | 110s | 19s | -83% |
| Node Variance | 4m 02s | 1m 07s | -72% |
| Compute/Run | ~43 min | ~34 min | -21% |
| Est. Annual Cost | $850 | $672 | -$178 |
Bonus: Test Code Improvements
Beyond workflow optimization, we also improved the test code itself by consolidating it blocks using RSpec's .and matcher chains. This reduced redundant before block executions.
For more on RSpec optimization techniques, see the RSpec official documentation on aggregate_failures and this practical guide on aggregating expectations.
Key Takeaways
Measure before optimizing
Understanding where time is spent (DB setup, test execution, node imbalance) guides which optimizations will have the most impact.
Runtime-based splitting is worth the setup
The caching mechanism adds some complexity, but the improvement in test distribution is significant.
Question default configurations
The redundant DB setup was likely copied from another context. Always verify that CI configurations match your actual requirements.