I once let Claude Code build and install my own iOS app. It reported “build succeeded” and “installed” again and again, yet not one of the fixes from that afternoon had reached the device.
The short answer: the cause was shell behavior, not the AI making things up. When you pipe a command into tail, as in flutter build ios ... 2>&1 | tail -60, the exit status comes from tail. A failed build still returns 0, so nobody notices. The same kind of accident happened again with zsh’s noclobber option.
This article covers both incidents, how I reproduced them in zsh 5.9 and bash 3.2.57, and the rules I now give the AI.
Sponsored
Why did the AI report a failed build as a success?
The build command was piped into tail, so the exit status was 0 even though the build failed.
On July 26, 2026, I had Claude Code repeatedly fix and rebuild a Flutter app. Every report said the build succeeded and the app was installed. But on the device, none of the afternoon’s changes were there. The installed app was still the build from 12:09 that day.
It turned out the AI verified each change with a command like this, which keeps only the last 60 lines of a long log:
flutter build ios --release 2>&1 | tail -60
In reality, every build had failed. The culprit was the Swift code for the home screen widget, which declared a function inside a ViewBuilder closure.
Swift Compiler Error: Closure containing a declaration cannot be used with result builder 'ViewBuilder'
The compile error was not the real problem. The real problem was that a failed build returned exit status 0, so the failure was invisible. Several rounds of changes were stacked on top of a build that never succeeded. I collected the build errors that widget extensions tend to cause in Flutter Widget Extension Build Errors: 3 Fixes for Xcode.
Which command’s exit status does a pipeline return?
A pipeline returns the exit status of its last command. If an earlier command fails but the last one succeeds, the result is 0.
I checked this in zsh 5.9. The false command always fails with exit status 1.
% false | tail -1; echo "exit=$?"
exit=0
macOS’s default bash 3.2.57 also printed exit=0. Both shells provide a way to read the status of each command in the pipeline and a way to make any failure fail the whole pipeline.
| Shell | Read an earlier command’s status | Fail the pipeline if any command fails |
|---|---|---|
| zsh 5.9 | ${pipestatus[1]} (1-based array) |
setopt pipefail |
| bash 3.2.57 | ${PIPESTATUS[0]} (0-based array) |
set -o pipefail |
In zsh, pipestatus still held the 1 from false, and with pipefail enabled the pipeline itself returned 1:
% false | tail -1; echo "pipestatus=${pipestatus[*]}"
pipestatus=1 0
% setopt pipefail
% false | tail -1; echo "exit=$?"
exit=1
Why did a failed file write with noclobber still look “done”?
With zsh’s noclobber enabled, redirecting with > to an existing file fails. But the next line of the script still runs, so if the last command succeeds, the whole thing looks successful.
On September 7, 2026, I had the AI write a code audit to IMPROVEMENTS.md. It created the file with cat > IMPROVEMENTS.md <<EOF and reported the task as done. The file still contained the previous audit, without a single line of the new results.
noclobber is an option that prevents > from accidentally overwriting files. It was enabled in my zsh, and the shell the AI used picked up that setting too.
Here is the reproduction in zsh 5.9:
% echo old > nc.txt
% setopt noclobber
% echo new > nc.txt
zsh: file exists: nc.txt
% echo "exit=$?"
exit=1
% cat nc.txt
old
The redirect itself fails with status 1. The trouble is that the following lines keep running as if nothing happened. A final echo "Created" succeeds, and the script as a whole looks fine.
bash 3.2.57 behaves the same with set -o noclobber, and prints bash: nc2.txt: cannot overwrite existing file. When you really do want to overwrite, >| works even with noclobber enabled.
% echo new >| nc.txt
% cat nc.txt
new
Sponsored
How should you verify an AI’s “success” report?
Check the exit status without altering it, and check the produced artifact itself.
After these two incidents, I added the following rules to the instructions I give the AI (Claude Code’s CLAUDE.md and command definitions).
▼Rules the AI follows now
① Never pipe build, test or analysis output. Write the log to a file and print the exit status right after
② Before reporting success, check that the artifact is newer than the source
③ After generating a file, read it back (use >| when overwriting is intended)
Rule 1 looks like this. If you still want to read only the end of the log, read the file after capturing the exit status.
flutter build ios --release > build.log 2>&1
echo "exit status: $?"
tail -60 build.log
How do you check that the artifact is newer than the source?
The test command’s -nt (newer than) operator compares two files’ modification times.
This check would have caught the July 26 incident. The app binary on the device was built at 12:09, which was older than the source files edited that afternoon.
if [ build/ios/iphoneos/Runner.app/Runner -nt lib/main.dart ]; then
echo "artifact is newer than the source"
else
echo "artifact is stale (the build did not take effect)"
fi
For apps with extension targets, such as widgets, compare the extension binaries inside Runner.app/PlugIns/ as well. The app binary can be fresh while the extension build failed.
Why do AI agents fall into this pattern so easily?
Build and test logs often run to hundreds of lines, and AI agents frequently trim them with tail or grep to read less.
In the work I delegate, this pattern comes up again and again. Reading only the end of a log is reasonable by itself. The problem is that trimming the output for readability also erases the evidence of failure.
A report that does not mention failure does not prove that nothing failed. In Should Test Cases Include Steps? I wrote about how checks that were never performed simply do not appear in a report. AI reports have the same structure.
For another example of controlling what an AI does before and after its work, I used a Claude Code hook to clean up leftover processes in Playwright MCP: Fixing the Stuck Browser Reuse Problem.
Summary
The AI reported failed builds as successes because of two shell behaviors:
▼Key points
① A pipeline returns the exit status of its last command. Piping into tail hides a failed build
② Earlier failures are visible in pipestatus (zsh) and PIPESTATUS (bash), and pipefail makes the pipeline fail
③ With noclobber, > to an existing file fails, but the next line still runs
④ Verify an AI’s “success” by logging to a file and checking the exit status, and by checking that the artifact is new
The most painful part was building half a day of work on top of a failed build. Rather than distrusting the AI’s report, I now first question whether the exit status behind that report was captured correctly. Since adopting these rules, I have not had the same kind of accident again.
Tested with zsh 5.9 and GNU bash 3.2.57 on macOS on September 12, 2026.