When parallel tests disturbed each other through global notifications
DailySudoku, a Sudoku app I build as a side project, gives one free game a day and shows a banner ad. It is available on the App Store, Google Play, and the web. A failure showed up in the iOS tests for that banner ad.
A test named holdsImpressionUntilTheSlotIsLiveAgain passed locally in about 0.02 seconds, while the CI job ran into its 45-minute limit. After the first fix the endless wait was gone, but the test kept failing. The code baseline is the iOS implementation at commits d44808bb and b7646ee2 on the develop branch as of August 14, 2026.
The notifications the banner listened to
AdBannerView has to react to a remove-ads purchase, a consent change, and a return to the foreground. So it observed three notifications on NotificationCenter.default.
This is the code where the banner registered its observers.
for name in [SudokuStore.didChange, AdsNotifications.didChange, UIApplication.didBecomeActiveNotification] {
NotificationCenter.default.addObserver(self, selector: #selector(refresh), name: name, object: nil)
}
The last one, the did-become-active notification, is there for a reason. Something has to restart the load and the impression handling that were held back while the slot was inactive. When the user returns to a screen that is already showing, viewWillAppear cannot be expected to be called. Without this trigger, both the ad slot and the impression record can stay empty.
At first I suspected load on the runner
Several jobs share the CI runner, so load was my first suspect. Then I found a different problem in the test code. Execution continued after #expect failed, and it then entered a while with no upper bound that waited for shownKind.
Reduced to its shape, the flow looked like this. The arguments and the loop body are omitted.
#expect(await waitForChain( // 실패해도 실행이 멈추지 않는다
// …
))
while banner.shownKind == nil { // 그래서 여기서 영원히 대기한다
// …
}
The Korean comment on the first line says that execution does not stop even on failure, and the one on the while line says that it therefore waits here forever.
I measured the time too. Run alone, this test took 0.024 seconds locally. Under contention, there was a case where the asynchronous MainActor work that builds the chain was delayed beyond the 4-second observation window.
I also ruled out other changes as the cause. releaseForcedProvider in the PR that changed ad routing is always nil in DEBUG. The PR under review has no change in the ad directory. Both were confirmed from the ancestry and the diff of the two PRs.
Defensive changes went in first. The wait was raised from 4 seconds to 8 to match the provider’s network time limit, and waitForShown was built on the existing polling approach. #require stops progress after a failure, and the unbounded loop became a wait with a time limit. The result was that a 45-minute wait turned into an 8-second failure. The interference between tests was still there.
A notification from another test reset the banner
The hypothesis is the combination of parallel execution in Swift Testing and a notification center shared by the whole process. A static search found 7 test files that post SudokuStore.didChange. A notification posted by an unrelated test synchronously calls refresh() on the banners that are observing. When ads are disabled, the !enabled branch clears chain, shownKind, and pendingImpression and bumps the generation.
Two conditions made only this test vulnerable.
- The watcher uses
shownKindas input to decide the remove-ads state. A reset can happen once the value exists. A notification before the value exists only callssetNeedsLayout(). - Unlike the other tests, this one does not call
layoutIfNeeded()before its first wait. So after a reset there is no layout pass to restart the load.
I judged that this is the only test where both conditions hold together.
The same hypothesis explained the difference between local runs and CI. When the test passed locally, the suites running alongside it presumably did not post a notification during the risky window. The observation that contention delays timing was valid in itself, but it did not fit as the cause of this failure.
I confirmed the hypothesis with an experiment. Adding a Task that posts the notification made the failure reproducible locally. The existing test, run in parallel, failed at the same place in the code.
Injecting the notification center
I added notificationCenter: NotificationCenter = .default to the initializer of AdBannerView and changed the observers to register on the injected center.
This is part of the changed initializer. The setup code before the observer registration is omitted.
init(
colors: SudokuColors,
adsProvider: AdsProvider,
consentProvider: ConsentProvider,
notificationCenter: NotificationCenter = .default // 추가
) {
// …
for name in [SudokuStore.didChange, AdsNotifications.didChange, UIApplication.didBecomeActiveNotification] {
notificationCenter.addObserver(self, selector: #selector(refresh), name: name, object: nil)
}
}
The Korean comment next to the new parameter means “added.”
Three test files each pass their own instance. Because of the default value, the app keeps using the same NotificationCenter as before. adsProvider, consentProvider, makeAd, and logger were already injected, so this follows the same pattern. The reset behavior in refresh() and didMoveToWindow() means that the creative is removed, so I left it as it was.
Other options were considered. @Suite(.serialized) only limits concurrency inside that suite, so it has no effect on notifications posted by other suites. And when layout does not run again, raising 8 seconds to 80, or waiting forever, does not recover the test.
How the fix was confirmed
xcodebuild test reported 299 tests in 60 suites passing. One regression test was added, which took the count from 298 to 299. The problem test alone ran in 0.023 seconds, and the whole run took 8.58 seconds.
I also checked that the new test catches the defect. With a mutation that puts the center back to .default, both the new test and the existing test failed. Deliberately putting a failing condition into verification code is covered in the post on applying a negative control to a verification script.
Four rules came out of this. A failure that happens only in CI is not assumed to be a performance problem. A hypothesis built from reading code is confirmed by reproduction. A regression test is checked by putting the defect back and seeing it fail. A time limit and a stop after failure have value apart from the isolation fix, so they stay.
What I observed directly is the failure under an artificially created notification. I did not find which suite actually posted the notification in the CI run that failed. The search result of 7 files cannot settle who the sender was at that time. Whether the centers in the three test files are also independent per test case is not covered in this post.
이 포스팅은 쿠팡 파트너스 활동의 일환으로, 이에 따른 일정액의 수수료를 제공받습니다.