Skip to content

[codex] parallelize release code generation - #27702

Merged
tamird merged 1 commit into
mainfrom
optimize-release-codegen-units
Jun 12, 2026
Merged

[codex] parallelize release code generation#27702
tamird merged 1 commit into
mainfrom
optimize-release-codegen-units

Conversation

@tamird

@tamird tamird commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

The release profile still uses one codegen unit, which serializes LLVM code generation within each crate. That setting was selected alongside fat LTO for optimization quality and binary size, but releases now use ThinLTO and code generation dominates the critical-path build.

Use four codegen units. On an Apple M4 Max with 16 cores and 128 GiB RAM, using rustc 1.96.0, four and eight units took 507.486 and 505.325 seconds respectively. Four therefore keeps the build-time gain while limiting the stripped codex increase to 14.7%, compared with 21.5% at eight units. The gzip-compressed binary grows 7.8% at four units.

The one-unit build from an empty target directory took 981.150 seconds. That comparison also populated dependency and native build caches, so it is directional rather than controlled. It agrees with the earlier clean matrix where eight units reduced 671 seconds to 303 seconds: https://gist.github.com/anp/4b88393a0acd35783d9f42156f3243d5

At the local 48% reduction, the current release's 55m22s critical-path macOS Cargo step would save about 26 minutes from the 71m28s workflow: https://github.com/openai/codex/actions/runs/27367405663

The prompt-image medians ranged from 3.9% faster to 0.9% slower. CLI startup shifted by 1-2 ms while user and system CPU time were unchanged.

This is a draft because the release-latency improvement may not justify the binary-size increase.

The release profile still serializes code generation within each crate.
That setting was introduced with fat LTO to improve optimization and
binary size. Releases now use ThinLTO, and code generation dominates the
critical-path build.

Use four codegen units. On an Apple M4 Max with 16 cores and 128 GiB
RAM, using rustc 1.96.0, four and eight units took 507.486 and 505.325
seconds respectively. Four keeps the build-time gain while limiting the
stripped codex increase to 14.7%, compared with 21.5% at eight units.
The gzip-compressed binary grows only 7.8% at four units.

The one-unit build from an empty target directory took 981.150 seconds.
That comparison also populated dependency and native build caches, so it
is only directional. It agrees with an earlier clean matrix where eight
units reduced 671 seconds to 303 seconds:
https://gist.github.com/anp/4b88393a0acd35783d9f42156f3243d5

At the local 48% reduction, the current release's 55m22s critical-path
macOS Cargo step would save about 26 minutes from the 71m28s workflow:
https://github.com/openai/codex/actions/runs/27367405663

The prompt-image medians ranged from 3.9% faster to 0.9% slower. CLI
startup shifted by 1-2 ms while user and system CPU time were unchanged.
@tamird
tamird marked this pull request as ready for review June 12, 2026 02:44
@tamird
tamird merged commit e23d4df into main Jun 12, 2026
46 checks passed
@tamird
tamird deleted the optimize-release-codegen-units branch June 12, 2026 02:44
@github-actions github-actions Bot locked and limited conversation to collaborators Jun 12, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants