iOS Tooling Should Keep Up with Coding Agents
The iOS community deserves tooling designed for how we increasingly work: several coding agents editing, generating, linting, building, testing, and merging throughout the day. In that workflow, “written in Swift” tells me very little about whether a tool is a good choice. I care about the correctness of its output and how quickly I can get the next useful result.
I love Tuist. Its module cache is a substantial part of what makes my modular app practical to work on, and I appreciate how quickly the project evolves. That experience is exactly why I wanted to investigate the remaining cost of workspace generation. Good tooling gives us a foundation to improve further.
Workspace Generation in Rust
I first asked a coding agent to implement the Tuist generation workflows I use in Rust: app, explicit targets, changed modules, and all local modules. The prototype uses a native JSON description instead of Swift manifests and writes Xcode projects and a workspace. It reuses prepared binary artifacts from Tuist; dependency resolution and remote caching remain outside its scope.
These are generation medians from five samples per implementation and scenario on October 8, using alternating serial runs on a shared Apple Silicon workstation. Installation, cache production, and app compilation are excluded.
| Cache state | Tuist generation | Rust prototype | Difference |
|---|---|---|---|
| Warm | 7.58 s | 3.01 s | 4.57 s |
| Implementation edited | 15.84 s | 6.49 s | 9.35 s |
| Interface edited | 15.00 s | 2.63 s | 12.37 s |
| Source file added | 12.70 s | 2.28 s | 10.42 s |
I compared the generated projects and workspace, including targets, sources, resources, dependencies, selected settings, and scheme actions. Both apps built in all four scenarios, with the changed files actually compiled. Relevant tests passed separately. This is a limited prototype with a different manifest format, so the result does not isolate the effect of the language or establish full Tuist compatibility.
The Same Lint Rules in Four Languages
I also had the agent implement our custom lint rules in Rust, Swift, Ruby, and Python. Every implementation had to produce matching findings. The comparison used 40 randomized blocks per workload, warm caches, and precompiled optimized Rust and Swift binaries.
| Implementation | Twenty-file feedback proxy | Full-repository lint workflow |
|---|---|---|
| Rust | 0.205 s | 1.579 s |
| Ruby | 0.322 s | 2.842 s |
| Python | 0.296 s | 3.213 s |
| Swift | 0.709 s | 3.773 s |
The repository contained 1,153 tracked Swift files. Both workflows included the same upstream SwiftLint invocation; the full workflow also included module checks. Compilation was excluded. The small-file proxy shows a modest saving versus Python; full-repository checks show a larger one. These results compare particular implementations, including different algorithms and regex engines. They give us concrete reasons to question our defaults.
What Repeats Across 100 Threads
Across the same 100 Codex threads, I counted approximately 1,424 generation requests and 870 lint requests. Applying the measured medians gives these combined savings from moving Tuist generation and our Python lint rules to Rust:
| Generation scenario | Combined saving across 100 threads | Average saving per thread |
|---|---|---|
| Warm cache | 1 h 51 min | 1 min 7 s |
| Implementation edited | 3 h 44 min | 2 min 15 s |
| Interface edited | 4 h 56 min | 2 min 58 s |
| Source file added | 4 h 10 min | 2 min 30 s |
Each row assumes all generation requests use that cache scenario. Lint adds approximately 2 min 35 s to each total. These are modeled tool-time savings across all 100 threads; the actual cache-state mix and production wrapper overhead remain unmeasured.
Sometimes I start a thread and return much later, so finishing sooner makes no difference to me. But when I am actively steering an agent, waiting for the next build or test result, those repeated seconds are directly in the way. That is the feedback loop I want our tooling to serve.
Stop Selling Familiar Syntax
I am a solo developer. These differences are already visible in one app on one workstation. A team of twelve or thirty developers has more work moving through the same kinds of checks. Its savings need its own measurements, but the priority should be obvious: make the machinery keep up with the people and agents using it.
Instead, too much of the pitch for Swift tooling still comes back to how nice it feels. Familiar syntax. Autocomplete in Xcode. An API that looks like the code in your app. I think we should stop treating those as decisive arguments.
Autocomplete is not a reason to wait longer for the same result.
Tuist’s manifest documentation describes compiler validation, code reuse, and Xcode’s editing support. Those benefits made sense when manually writing and editing manifests was a large part of the experience. They carry much less weight in my current workflow. An agent writes the plumbing. I steer it, inspect the changes that matter, and validate what comes out.
I am not reading every implementation behind every command I run. I am certainly not choosing them for the pleasure of reading their source. I want the right projects, the right findings, and the next test result. Whether the engine underneath uses Swift, Rust, or another language is secondary to delivering those outcomes quickly and reliably.
Type safety is useful. It is also available outside Swift. Familiarity is useful. It is also something an agent can help us work past. Neither should give a slower implementation a free pass. Show me the output and the timings. A nicer editing experience does not settle that comparison.
Your Own Tools Are Fine
I already have tooling calibrated to my app: lint rules for our architecture, target-selection scripts, worktree helpers, and validation commands. They exist to make this particular workflow work. They do not need to become general-purpose products before they are allowed to be useful.
That is an opportunity coding agents give us. We can ask for a small tool, define what it must produce, check it against real examples, and keep improving it. We can use a language we would not have chosen to maintain every line by hand. The result can be a simple executable that does its job and gets out of the way.
Call it vibe coding if you like. I am comfortable using it for local tooling with a clear scope and results I can verify. A malformed workspace or a broken helper is something I can inspect, regenerate, and fix. Iterating once more is an acceptable cost of improving my development environment. I do not need to treat every helper as though it were a payment flow shipping to customers.
The distinction is scope. Local helpers stay in the development environment; release automation and checks that can silently accept a broken app need stronger validation. I can accept rough edges in a helper while keeping the quality bar for the app. Those are separate decisions.
We should stop selling “written in Swift” as a quality label for iOS tooling.
Build the tool in whatever technology gives us correct results with the shortest useful feedback loop. Let agents handle unfamiliar implementation details. Let engineers define the behavior, measure it, and verify it. Personal tools are fine. Unfamiliar languages are fine. A tool earns its place by working well.
I want less celebration of how nice our tooling is to write, and more ambition about how fast it lets us work. The iOS community deserves that. Familiar syntax is no excuse for leaving a measured improvement on the table.