Enter the passphrase to continue.
First Student and Thoughtworks are building a new AI-powered software delivery lifecycle: prove it first inside First Alt, then roll it out across the company. This cockpit tracks that 3-year program live: how the new model compares to today's delivery, what it costs, and what we're learning. Thoughtworks brings modules from its Agentic Development Platform, AI/works™, to speed the work. Red vs Blue, the Phase 1 proof experiment, is running now.
3-year AI-powered SDLC transformation. The Red vs Blue proof experiment is the current early phase, running at Sprint 2.
Sprints 0–1 are actual, Sprint 2 is in progress (today). Sprints 3–6 are an illustrative projection, not a forecast we're committing to.
7.4% of the annual envelope spent through Day 27 of Year 1 ($15.0M across the 3-year program).
Every team moves the same way: Blue to Purple to Red, converting onto the new AI-powered SDLC. Today, the vanguard is the only team in motion, proving the new model out as the fleet's first Red team. Purple marks the next wave in line, queued to begin once the vanguard proves out. This rolls across First Alt first, then across all of First Student, over the 3-year program.
The same five steps for every wave, whether it is the 1st team through or the 12th.
At any moment the fleet is exactly three buckets: Red, Purple, Blue. Today only the vanguard, the fleet's first Red team, is converting; Purple holds the next wave in line, not yet started. As the vanguard proves the model out, Purple waves begin converting behind it, Red grows, and Blue shrinks. It is early: no wave has finished converting yet.
Actual through Sprint 2 (today, Jul 27, 2026). Sprints 3–6 are an illustrative projection, not a forecast we are committing to.
All 12 teams in the fleet, current state, and the wave each one is converting in or scheduled to convert in.
| Team | State | Wave | Size |
|---|---|---|---|
| Routing Engine | Red | Sprint 1 | 7 |
| Driver App | Purple | Not started, next up | 8 |
| Scheduling Core | Blue | Scheduled Sprint 3 | 6 |
| Parent Portal | Blue | Scheduled Sprint 4 | 7 |
| Fleet Telematics | Blue | Scheduled Sprint 5 | 9 |
| Trip Planning | Blue | Scheduled Sprint 6 | 6 |
| Compliance & Safety | Blue | Queued | 5 |
| Billing & Invoicing | Blue | Queued | 8 |
| Depot Operations | Blue | Queued | 7 |
| Reporting & Analytics | Blue | Queued | 6 |
| Onboarding Platform | Blue | Queued | 5 |
| Integrations | Blue | Queued | 9 |
More delivered, with fewer people doing it. This is the point of the conversion, and it moves with every wave.
Story points delivered per sprint, not indexed. Actual through Sprint 2 (today); Sprints 3–6 are an illustrative projection.
The Red team is proving out the new AI-powered SDLC. The Blue teams are running First Student's current, traditional SDLC unchanged. Comparing them head-to-head is how we prove the new model before rolling it out further.
Actual data through Sprint 2 (today, in progress). Sprints 3–6 are an illustrative projection, not a forecast we're committing to.
| Red Team | Blue Teams | Delta | |
|---|---|---|---|
| Delivery model | Disruption model | Traditional ways of working | N/A |
| Tooling | AI/works™ modules (2 live, 2 building/exploring) | Existing toolchain | N/A |
| Status | Active | Running today's SDLC, unchanged | N/A |
| Sprint 2 snapshot | |||
| Team size | 7 people | 9 people | -2 ·-22% |
| Story points / sprint | 61 | 45 | +16 ·+36% |
| Velocity (3-sprint avg) | 52 | 44 | +8 ·+18% |
| Cycle time (median, days) | 5.2 | 6.5 | -1.3d ·-20% |
| Deploy frequency (per week) | 1.8 | 1.2 | +0.6/wk ·+50% |
| Change-fail rate | 15% | 18% | -3pts ·-17% |
Three more delivery signals point the same direction as the story-point chart above.
Cycle time and change-fail rate are inverted on the y-axis so improvement always trends up, matching the chart above. Actual values are labeled at each end.
The toolset behind the new SDLC we're building at First Alt. Seven modules, one decision each: install it, build it out, explore it, or leave it out for now.
Connector shows whether work is actively flowing into the next module: solid and moving when it is, dashed when partial, dotted when the upstream module has nothing to send or the next module isn't consuming it yet.
| Module | Status | Owner | Teams using | Adopted | Activity this sprint |
|---|---|---|---|---|---|
| Reverse Engineering | Live | TW delivery lead | 2 of 12 | Jul 2, 2026 | 180 runs |
| Knowledge Fabric | Building | TW knowledge engineering lead | 1 of 12 | Since Jul 16 | 85 docs indexed |
| Dynamic Spec Generation | Not started | n/a | 0 of 12 | – | – |
| Rapid Prototyping | Live | TW product engineering lead | 2 of 12 | Jul 8, 2026 | 140 prototypes |
| Spec to Code | Not in use | n/a | 0 of 12 | – | – |
| Verification | Building | TW quality engineering lead | 2 of 12 | Since Jul 19 | 210 checks |
| Evals | Not started | n/a | 0 of 12 | – | – |
Actual through Sprint 2, today. Sprint 3–6 are an illustrative projection, not a commitment. Spec to Code and Evals are deliberate scope decisions, not delays. See Module status above.
The new AI-powered SDLC we're building at First Alt, with AI/works™ as the toolset inside it. Every module gets more autonomous as it proves itself: this is where that shows up, in capability maturity, the agents running today, and how the platform is compounding sprint over sprint.
Autonomy stage per module, Sprint 2 of the 3‑year program (Day 27 of 1,096).
| Module | Manual | Assisted | Semi‑autonomous | Autonomous | Last advanced |
|---|---|---|---|---|---|
|
Reverse Engineering
Ingests legacy First Alt code, reconstructs as‑is specs
Live
|
Sprint 2 | ||||
|
Knowledge Fabric
Loads and indexes First Alt specs, standards, and context
Building
|
Since kickoff | ||||
|
Dynamic Spec Generation
Drafts candidate specs from ticket descriptions and prior art
Not started
|
Not yet | ||||
|
Rapid Prototyping
Turns ideas into disposable prototypes ahead of a full spec
Live
|
Sprint 1 | ||||
|
Spec to Code
Generates implementation code directly from an approved spec
Not started
|
Not yet | ||||
|
Verification
Scores generated code against spec and quality bar
Building
|
Since kickoff | ||||
|
Evals
Runs the held‑out test suite and scores output quality
Not started
|
Not yet |
Every agent running against First Alt work today, sortable on any column.
| Agent | Purpose | Used by | Status | Invocations, Sprint 2 | Adopted |
|---|---|---|---|---|---|
| Claude Code | Generates code directly from spec, doing the job the Spec to Code module would otherwise do | Red team | Live | 96 | Jul 2, 2026 |
| Reverse engineering agent | Ingests legacy First Alt code and reconstructs as-is specs | Red team | Live | 54 | Jul 3, 2026 |
| Rapid prototyping agent | Turns ideas into disposable prototypes ahead of a full spec | Red team | Live | 41 | Jul 9, 2026 |
| Knowledge fabric agent | Loads and indexes First Alt specs, standards, and context for every other agent to draw from | Red team | Building | 28 | Jul 16, 2026 |
| Verification agent | Scores generated code against spec and quality bar, run on Red team code so far | Red team | Building | 19 | Jul 19, 2026 |
| Story point estimator | Suggests point estimates from historical velocity and spec complexity | Red team | Live | 12 | Jul 21, 2026 |
| PR summary agent | Drafts pull request descriptions and reviewer checklists from the diff | Red team | Building | 6 | Jul 24, 2026 |
| Regression triage agent | Clusters failing test output and proposes the likely root cause | Purple team | Not started | 0 | Not yet |
| Dynamic spec generation agent | Drafts candidate specs from ticket descriptions and prior art for review | Purple team | Not started | 0 | Not yet |
| Eval harness agent | Runs the module's held-out test suite and scores output quality | Red team | Not started | 0 | Not yet |
No agents match “”.
The platform compounds sprint over sprint: more live modules, more teams drawing on them, more work reused instead of rebuilt.
Actual through Sprint 2, today; Sprint 3–4 are an illustrative projection, not a commitment. Capabilities and teams onboarded are shown as a percent of program totals (7 modules, 12 fleet teams) so all three series read on one axis. Actual counts shown on hover.
The near-term queue for moving modules and agents forward.
Real delivery and platform work from Red and Purple teams. Filter by team, module, sprint, or status to find what you're looking for.
Full Red team velocity across all workstreams. Sprint 2 is in progress, so its bar reflects points shipped so far, not a final total. The item lists below include both Red and Purple teams and are a representative sample, not the complete set.
| ID | Title | Module | Team | Sprint | Status | Pts |
|---|---|---|---|---|---|---|
| FA-107 | Load First Alt design specs into the Knowledge Fabric | Knowledge Fabric | Purple Team 1 | Sprint 2 | In Progress | 5 |
| FA-108 | Wire verification scoring into CI/CD for Red team code | CI/CD & Tooling | Red Team 1 | Sprint 2 | In Progress | 8 |
| FA-109 | Retrain the route optimization model on First Alt lane data | Route Optimization | Red Team 2 | Sprint 2 | In Progress | 13 |
| FA-110 | Driver app offline-mode sync for rural routes | Driver App | Red Team 3 | Sprint 2 | In Review | 5 |
| FA-111 | Onboard Purple Team 1 with Red Team 1 mentors | Baseline & ROI | Purple Team 1 | Sprint 2 | In Progress | 3 |
| FA-112 | Stand up financial ROI tracking against the Sprint 0 baseline | Baseline & ROI | Red Team 1 | Sprint 2 | Blocked | 3 |
| No in-flight items match the current filters. | ||||||
| ID | Title | Module | Team | Sprint | Completed | Pts | Demo |
|---|---|---|---|---|---|---|---|
| FA-101 | Capture the Blue team's 12-month baseline | Baseline & ROI | Red Team 1 | Sprint 1 | Jul 3, 2026 | 8 | View demo |
| FA-102 | Reverse-engineer the legacy routing engine | Reverse Engineering | Red Team 1 | Sprint 1 | Jul 6, 2026 | 13 | View demo |
| FA-103 | Provision AI/works™ module access for the Red team | CI/CD & Tooling | Red Team 1 | Sprint 1 | Jul 7, 2026 | 3 | View demo |
| FA-104 | Knowledge Fabric ingestion of routing docs | Knowledge Fabric | Red Team 1 | Sprint 1 | Jul 10, 2026 | 5 | View demo |
| FA-105 | Prototype driver app trip-status notifications | Rapid Prototyping | Red Team 2 | Sprint 1 | Jul 12, 2026 | 5 | View demo |
| FA-106 | Add verification scoring to Red Team 1's PR checks | CI/CD & Tooling | Red Team 1 | Sprint 1 | Jul 14, 2026 | 8 | View demo |
| No completed items match the current filters. | |||||||
Sprint 2, Day 27 of 1,096. These metrics prove whether the new AI-powered SDLC outperforms the current SDLC, measured against the 12-month pre-program traditional-SDLC baseline captured before kickoff on Jul 1, 2026.
A Red Team vs. Blue Team experiment with a 1-year historical baseline is the gold-standard approach for evaluating AI in software engineering. The finding this framework is built around: AI acts as an amplifier. It accelerates code generation, but without balanced metrics that extra velocity spills into downstream bottlenecks: bloated PRs, code-review fatigue, CI/CD queue delays. These 9 metrics give a 360-degree view of speed, quality, cost, and developer experience across DORA, SPACE, and DevEx.
Do not measure lines of code, commit counts, or AI token usage as productivity metrics. AI makes it trivial to inflate commit volume and LOC, which incentivizes verbose, bloated architectures instead of elegant, maintainable code. Measure system outcomes and developer experience, not raw output.
| Framework | Recommended Metric | Dimension Tracked | Primary Question Answered |
|---|---|---|---|
| DORA | Lead Time for Changes | Delivery Speed | Does AI code reach users faster end to end? |
| DORA | Deployment Frequency | Batch Size / Velocity | Is the team shipping small, controllable increments? |
| DORA | Change Failure Rate | Stability | Does AI introduce hidden production defects? |
| SPACE | PR Review Turnaround Time | Collaboration | Are reviewers overloaded by AI-generated PRs? |
| SPACE | Cost per Delivered Value Unit | Performance / ROI | Is the AI tool cost justified by output gains? |
| SPACE DEVEX | Work Restart / Rework Rate | Efficiency / Quality | How much AI code gets rejected in review? |
| DEVEX | Perceived Cognitive Load | Mental Toil | Does AI reduce tedious tasks or add review stress? |
| DEVEX | Uninterrupted Flow Time | Flow State | Does AI support deep engineering work? |
| DEVEX | CI/CD Pipeline Duration | Feedback Loops | Is automation keeping up with code generation? |
Speed. Lead Time for Changes and Deployment Frequency, Sprints 0–2 actual, projected through Sprint 6. Open a tile for what it measures and why it matters.
Quality. Change Failure Rate and Work Restart / Rework Rate, Sprints 0–2 actual, projected through Sprint 6.
Cost and efficiency. Cost per Delivered Value Unit and PR Review Turnaround Time, Sprints 0–2 actual, projected through Sprint 6.
DevEx and SPACE. Perceived Cognitive Load, Uninterrupted Flow Time, and CI/CD Pipeline Duration, Sprints 0–2 actual, projected through Sprint 6.
Story points completed per sprint, Red team vs Blue team. The output side of Deployment Frequency and Lead Time above.
$0.37M spent through Day 27 of the 3-year, $15.0M program building and proving the new AI-powered SDLC ($5.0M Year 1 envelope). As Blue teams begin converting to Red, the cost to deliver a story point is starting to decline sprint over sprint.
| Category | Spend | Share of total |
|---|---|---|
| Delivery | $255,300 | 69.0% |
| Platform & enablement | $55,130 | 14.9% |
| Tooling & licenses | $35,150 | 9.5% |
| Transition costs | $24,420 | 6.6% |
| Total spend to date | $370,000 | 100.0% |
Figures on this page are illustrative sample data for this step-1 prototype and are not yet wired to live financial systems. Sprints 3 through 6 are projections at the current run-rate trend, not a committed forecast.
A running log of what proving the new AI-powered SDLC is teaching us, sprint by sprint, and how those lessons get compounded into the standard playbook instead of just filed away.
Proof the program compounds what it learns, not just accumulates notes. A lesson only counts once it changes how the next team works.
Describe the program as proving a new AI-powered SDLC, not as "Red team on AI/works." AI/works is a tool inside the new process, not the headline.
Two steerco updates got contested when "AI/works" sounded like the goal instead of a toolset inside the new delivery lifecycle being proven. The thesis is the new SDLC; AI/works is one of the things running inside it.
Run the Verification module on Blue-team code too, not just Red, and score both against the same rubric.
Early comparisons scored Red-produced code through Verification but took Blue's own word for its quality. That isn't a fair comparison. Both paths now run the identical automated rubric before any Red-versus-Blue number ships to steerco.
Sequence onboarding as tooling access, then conventions, then first paired task under the new SDLC. The vanguard team lost two days when a paired task was scheduled before access cleared.
The team converting from Blue to Purple lost roughly two days waiting on credentials because its first paired task under the new SDLC was scheduled before tooling access cleared.
Capture the twelve-month baseline before any new-SDLC tooling touches the team, even a small pilot. One live module during baseline week invalidates the comparison.
One team ran Verification in read-only mode during what was meant to be a clean baseline week, and its throughput numbers had to be thrown out and re-captured a sprint late.
Flag a platform gap the day it blocks someone, not at the retro. A late escalation is a lost day on the new SDLC, not a lesson.
A missing sandbox permission sat unescalated for several days because it first surfaced as a retro note instead of a same-day flag under the new delivery process. The fix took twenty minutes once someone actually raised it.
From proving a new AI-powered SDLC at First Alt to an outcome-based partnership across First Student. Day 27 of 1,096, Sprint 2, Phase 1 (the Red/Blue proof experiment) in progress.
Actual through today (Day 27, Jul 27, 2026), inside Phase 1 of Year 1. Everything beyond Phase 1 is an illustrative projection of the 3-year arc, not a forecast we're committing to. First Student's fleet composition in Years 2–3 is a placeholder pending scope.
First Alt's fleet of 12 teams keeps converting through the rest of Year 1, one Purple wave at a time. First Student's much larger fleet joins the rolling model in Year 2.
| Milestone | Day | Date | Phase | Status | Team |
|---|---|---|---|---|---|
| Program kickoff & governance alignment | 1 | Jul 1, 2026 | Phase 1 | Done | Program team |
| Red vanguard onboarded, system access granted | 3 | Jul 3, 2026 | Phase 1 | Done | Red team |
| AI/works™ platform install begins | 6 | Jul 6, 2026 | Phase 1 | Done | Platform team |
| Blue team 12-month baseline locked | 18 | Jul 18, 2026 | Phase 1 | Done | Blue teams |
| Sprint 1 complete, Sprint 2 begins | 22 | Jul 22, 2026 | Phase 1 | Done | Red & Blue teams |
| Sprint 2 comparison check-in at steerco (today) | 27 | Jul 27, 2026 | Phase 1 | In progress | Red & Blue teams |
| Phase 1 proof read-out: Red vs. Blue result at First Alt | 90 | ~Sep 28, 2026 | Phase 1 | Upcoming | Program team |
| Second wave of First Alt teams begins converting | 180 | ~Dec 27, 2026 | Year 1 | Upcoming | Platform team |
| Year 1 complete: new SDLC proven at First Alt, fan-out across First Student begins | 365 | Jul 1, 2027 | Year 1 | Upcoming | Program team |
| Year 2 complete: new SDLC fanned out across First Student | 730 | Jul 1, 2028 | Year 2 | Upcoming | Program team |
| Year 3 complete: program scaled to an outcome-based partnership | 1,096 | Jul 1, 2029 | Year 3 | Upcoming | Program team |