Back to blog
CI/CD GitHub Actions Rails

How We Reduced CI Time by 33% with Runtime-Based Test Splitting

A practical guide to optimizing parallel test execution in GitHub Actions

5 min read

33%

Time Reduction

21%

Cost Savings

83%

DB Setup Faster

The Problem

Our CI pipeline was taking over 10 minutes to complete, with RSpec tests running across 7 parallel nodes. Despite the parallelization, we noticed significant imbalances between nodes - some finishing in 6 minutes while others took over 10 minutes.

The root causes were:

  • 1. Filesize-based splitting - Tests were distributed by file size, not execution time
  • 2. Redundant DB setup - Each node was creating 7 databases when only 1 was needed
  • 3. Over-provisioned nodes - More nodes than necessary for the workload

Before vs After Comparison

Before: 7 nodes
Before: 7 nodes with 4+ min variance

Slowest: 10m 12s · Variance: 4m+

After: 5 nodes
After: 5 nodes with 1 min variance

Slowest: 7m 24s · Variance: 1m

The Solution

1. Runtime-Based Splitting

Instead of distributing tests by file size, we switched to runtime-based splitting using parallel_tests (a gem for running tests in parallel across multiple CPU cores) and its runtime log feature.

How it works:

  1. 1. Each node records test execution times to a runtime log
  2. 2. Logs are uploaded as artifacts after each run
  3. 3. A merge job combines logs and caches the result
  4. 4. Next run uses cached logs for optimal distribution
- name: Run tests
  run: |
    RUNTIME_LOG="tmp/parallel_runtime_rspec.log"
    if [ -f "$RUNTIME_LOG" ]; then
      GROUP_BY="runtime"
    else
      GROUP_BY="filesize"
    fi
    bundle exec parallel_test spec -t rspec \
      --group-by "$GROUP_BY" \
      --runtime-log "$RUNTIME_LOG"

2. DB Setup Optimization

We discovered that each CI node was creating 7 databases during setup, even though each node has its own isolated MySQL service and only uses 1 database. Our project uses ridgepole (a DSL-based schema management tool) instead of Rails migrations.

Before (110s)

parallel:create[7] + ridgepole:parallel_prepare

After (19s)

db:create + ridgepole:apply

3. Node Count Reduction

With better test distribution and faster DB setup, we reduced from 7 to 5 nodes while maintaining acceptable wall time. This yielded a 21% reduction in compute costs.

Results

Metric Before After Change
Wall Time 10m 12s 6m 50s -33%
DB Setup 110s 19s -83%
Node Variance 4m 02s 1m 07s -72%
Compute/Run ~43 min ~34 min -21%
Est. Annual Cost $850 $672 -$178

Bonus: Test Code Improvements

Beyond workflow optimization, we also improved the test code itself by consolidating it blocks using RSpec's .and matcher chains. This reduced redundant before block executions.

For more on RSpec optimization techniques, see the RSpec official documentation on aggregate_failures and this practical guide on aggregating expectations.

Key Takeaways

Measure before optimizing

Understanding where time is spent (DB setup, test execution, node imbalance) guides which optimizations will have the most impact.

Runtime-based splitting is worth the setup

The caching mechanism adds some complexity, but the improvement in test distribution is significant.

Question default configurations

The redundant DB setup was likely copied from another context. Always verify that CI configurations match your actual requirements.

Tools Used