Turning Apple Pencil handwriting into Sudoku digits
The goal for Apple Pencil input in DailySudoku was simple. Writing a 7 on a cell places a 7, and a quick dot with the pencil selects that cell. When a palm touches the screen or the recognition result is uncertain, no digit should be placed.
Keeping that goal meant solving four things: receiving input without clashing with the existing board gestures, grouping several strokes into one character, converting the input into the same shape used in training, and deciding when to apply a result. Below is the implementation based on the develop branch as of 2026-08-07 (commit 6db30672), along with the parts that automated tests alone could not confirm.
Sending the pencil and the finger down different paths
I placed a child view, PencilDrawOverlay, on top of BoardView. PencilDrawOverlay is a transparent PKCanvasView with the .pencilOnly policy, so only the pencil can draw on it. The board’s recognizer, in turn, does not accept the pencil.
These are the touch types the board recognizer accepts.
허용: direct, indirect, indirectPointer
제외: pencil
In the block, 허용 means “allowed” and 제외 means “excluded.”
With this split, the pencil goes to the overlay, and the finger, mouse, and trackpad go to the original cell selection handling. While a stroke is being written, the board’s pan is turned off and tap stays on. The overlay is not exposed as an accessibility element. VoiceOver users keep using the 81 virtual cells described in the post about board accessibility. The cell geometry is shared, and only the path a touch takes is different.
Tap or ink, and which group
If the diagonal of the first stroke’s bounding box is 6pt or less and the stroke lasts 0.12 seconds or less, it is a tap. Both conditions must hold. For a tap, the ink is cleared and the cell at the center of the bounding box is selected. A stroke near a group that is already being written counts as ink even when it is small and short, because it can be part of a digit.
A group is finalized 0.45 seconds after the last stroke. Each added stroke resets the end of the wait. The value accounts for digits such as 4, 5, and 7 that can take more than one stroke.
If a stroke that is not a tap is more than 60pt away from the center of the current group in the x or y direction, the previous group is finalized and a new group starts. For a short, small stroke far away, the tap decision comes first. The previous group is processed and then the new cell is selected. The target cell comes from passing the center of the whole group’s box to BoardView.cellIndex. The rounded-corner check also uses the same function as the board. Classifying a stroke, assigning it to a group, and choosing the target cell are separate steps.
The weak point is that 60pt is a fixed distance. Cell spacing changes with screen size, and the boundary for writing in adjacent cells one after another is not settled by tests. I am considering a value proportional to the cell size, or measuring distances on supported devices and calibrating.
Building the 28×28 the model sees
Point lists are taken from the paths of PKDrawing and drawn again into a grayscale buffer the size of the board. Then the image is cropped and scaled to match the MNIST format.
This is the preprocessing flow in order.
PKDrawing
-> [[CGPoint]]
-> round-cap grayscale raster
-> ink bounding-box crop
-> longest side 20px, aspect ratio 유지
-> bilinear resample
-> center of mass를 28×28 중앙으로 이동
-> [1, 1, 28, 28] Float tensor
The Korean in the block says to keep the aspect ratio (유지) and to move the center of mass to the center of the 28×28 image.
The stroke width shown on screen is 6pt for normal input and 3pt for notes. For inference, both are redrawn at 6pt. This keeps the display mode from changing the distribution of the input the model receives.
The preprocessing tests check that empty input returns nil, that the same input gives the same result, that the longest side is 20px, and that the center of mass is in place. They also compare the Core ML label and probability from the Python reference with the Swift result. If any of the crop, scale, pixel direction, or intensity range differs from the training conditions, the model can run without an error and still misrecognize the digit. A successful run does not show that the preprocessing is correct.
Results to accept and results to drop
The model keeps all 10 classes, 0 through 9. Sudoku has no 0, but I kept it so that an ambiguous circle is not forced into 6 or 9. Input classified as 0 is dropped at the apply step.
The default confidence threshold is 0.8. For notes, where a correction costs less, it is lowered by 0.1 to 0.7. Only when the result is between 1 and 9 and meets the threshold does the app select the cell and place the digit. A 0, a value out of range, low confidence, a model loading failure, a preprocessing failure, and an inference error are all rejects. A reject does not increase the mistake count. It simply does nothing. The ink is cleared regardless of the result.
In the current structure, failures with different causes all merge into a single .reject. There is no way to tell whether the model failed to load, the probability was low, or the label was wrong. Counts per cause are needed. I am considering watching the failure rate with on-device statistics only, without sending the original strokes, but that is not implemented yet.
Results that arrive late
Rendering and prediction run in Task.detached, and an accepted result calls selectCell and placeDigit on the main actor. There are two problems here.
The first is sharing the model. CoreMLDigitRecognizer is declared @unchecked Sendable, so several tasks can use the same MLModel at once. The MLModel documentation sets a condition that calls on one instance be serialized. @unchecked Sendable only skips the compiler check and does not change the runtime condition of MLModel. So the current declaration may conflict with that condition.
This is a proposal to let an actor own the recognizer. It is not applied yet.
actor DigitInference {
private let recognizer: CoreMLDigitRecognizer
init() throws {
recognizer = try CoreMLDigitRecognizer()
}
func recognize(_ pixels: [Float]) throws -> DigitRecognition? {
try recognizer.recognize(pixels: pixels)
}
}
With this, predictions run one at a time, and callers have to use await. Cloning the model per queue is another option, but it is something to review only after measurements show that a single actor falls short on latency or throughput.
The second problem is that an old result can be applied after the state has changed. The completion checks only whether the game is in the playing state. After a pause and resume, or after switching note mode, a result requested earlier can still be applied. The two modes can also disagree: the threshold is chosen from the mode at capture time, while placeDigit acts on the mode at completion time. I am considering storing the three values below with each request and checking that they still match right before applying.
input revision + target cell + entry mode
If the values differ, the result is discarded. A monotonically increasing number works as the revision.
What automated tests covered and what they did not
The automated tests focus on the pure reducer and renderer. They cover hit testing on the rounded board, the tap and ink conditions, debounce, restart on a distant stroke, normalization, the reject branches, loading the bundled model, the Python and Swift comparison, the allowed touch list, and the delegate wiring.
There are only two handwriting fixtures, 1 and 7. That means only two kinds of input are checked from rendering through to the expected digit. The threshold search in the training script is also a proxy evaluation that uses MNIST with rotation and stroke-width changes, plus 0 and noise. So the recognition rate for real pencil handwriting, and whether 0.8 is the best value, are not proven.
For the delegate, there is a record in a source comment. It diagnosed that on iOS 26.5, PencilKit falls into recursion when the canvas sets itself as its own delegate, so a separate bridge is used. The test only checks whether the canvas and the delegate are the same object. I did not recheck the device logs from that time. Another limit is that the .pencilOnly path cannot be exercised with a mouse in the simulator.
Six things need to be checked on devices.
- Precision and reject rate for 1 through 9 across handwriting styles
- Latency of the first response and the order of consecutive results
- The 60pt boundary on each screen size
- Behavior when pencil, finger, pointer, and VoiceOver are mixed
- Discarding earlier results during pause and resume or a mode change
- Whether users can understand a model loading failure or repeated rejects
For metrics, I plan to look at the share of wrong inputs among accepted results, the reject rate, and p95 latency. My judgment was to protect the precision of applied actions first and accept that more inputs go unrecognized.
What it takes to switch to PKStrokeRecognizer in iOS 27
The iOS 27 beta has PKStrokeRecognizer. It is an actor, and it is an API that recognizes strokes as text asynchronously on the device. It takes a PKDrawing as input. Recognition results can differ by OS, so storing a result means keeping recognitionVersion with it. A guide to recognizing handwriting and converting it to text is published as well. An app with an iOS 26.5 deployment target needs the Xcode 27 beta SDK, #available(iOS 27.0, *), and a fallback path that uses Core ML.
Swapping only the implementation is hard. The current DigitRecognizing takes [Float] and returns a digit and a confidence. PKStrokeRecognizer takes input through updateDrawing(_:), returns String? from recognizedText(strokeIDs:), and has no confidence. So a higher asynchronous boundary that takes the drawing and its context is needed.
recognize(drawing, context) async -> candidate 또는 reject
Here 또는 means “or”: the call returns either a candidate or a reject.
The Core ML side can keep the current rendering and probability policy, and the new adapter can limit the text to 1 through 9. The rule for applying a result that has no probability automatically needs separate product validation. During the beta period I plan to keep the existing model.
The values 0.8, 0.7, and 60pt in this post are constants set by policy, not optimal values found by measurement. Two handwriting fixtures and a proxy evaluation on MNIST cannot speak for real accuracy, and the six items above stay open until they are checked on devices.
이 포스팅은 쿠팡 파트너스 활동의 일환으로, 이에 따른 일정액의 수수료를 제공받습니다.