# Turning Apple Pencil handwriting into Sudoku digits

> Placing a handwritten digit in DailySudoku: splitting strokes into taps and ink, grouping them, building model input, and applying only trusted results.

- Canonical: https://jaemyeong.com/en/blog/apple-pencil-handwriting-digit-input-pipeline/
- Published: 2026.08.07
- Updated: 2026.10.04
- Category: IT/개발
- Tags: #iOS, #Swift, #PencilKit, #Core ML, #Apple Pencil

The goal for Apple Pencil input in DailySudoku was simple. Writing a 7 on a cell places a 7, and a quick dot with the pencil selects that cell. When a palm touches the screen or the recognition result is uncertain, no digit should be placed.

Keeping that goal meant solving four things: receiving input without clashing with the existing board gestures, grouping several strokes into one character, converting the input into the same shape used in training, and deciding when to apply a result. Below is the implementation based on the develop branch as of 2026-08-07 (commit 6db30672), along with the parts that automated tests alone could not confirm.

## Sending the pencil and the finger down different paths

I placed a child view, `PencilDrawOverlay`, on top of `BoardView`. `PencilDrawOverlay` is a transparent `PKCanvasView` with the [`.pencilOnly` policy](https://developer.apple.com/documentation/pencilkit/pkcanvasviewdrawingpolicy), so only the pencil can draw on it. The board's recognizer, in turn, does not accept the pencil.

These are the touch types the board recognizer accepts.

```text
허용: direct, indirect, indirectPointer
제외: pencil
```

In the block, 허용 means "allowed" and 제외 means "excluded."

With this split, the pencil goes to the overlay, and the finger, mouse, and trackpad go to the original cell selection handling. While a stroke is being written, the board's pan is turned off and tap stays on. The overlay is not exposed as an accessibility element. VoiceOver users keep using the 81 virtual cells described in [the post about board accessibility](/en/posts/adaptive-accessible-sudoku-board/). The cell geometry is shared, and only the path a touch takes is different.

## Tap or ink, and which group

If the diagonal of the first stroke's bounding box is 6pt or less and the stroke lasts 0.12 seconds or less, it is a tap. Both conditions must hold. For a tap, the ink is cleared and the cell at the center of the bounding box is selected. A stroke near a group that is already being written counts as ink even when it is small and short, because it can be part of a digit.

A group is finalized 0.45 seconds after the last stroke. Each added stroke resets the end of the wait. The value accounts for digits such as 4, 5, and 7 that can take more than one stroke.

If a stroke that is not a tap is more than 60pt away from the center of the current group in the x or y direction, the previous group is finalized and a new group starts. For a short, small stroke far away, the tap decision comes first. The previous group is processed and then the new cell is selected. The target cell comes from passing the center of the whole group's box to `BoardView.cellIndex`. The rounded-corner check also uses the same function as the board. Classifying a stroke, assigning it to a group, and choosing the target cell are separate steps.

The weak point is that 60pt is a fixed distance. Cell spacing changes with screen size, and the boundary for writing in adjacent cells one after another is not settled by tests. I am considering a value proportional to the cell size, or measuring distances on supported devices and calibrating.

## Building the 28×28 the model sees

Point lists are taken from the paths of `PKDrawing` and drawn again into a grayscale buffer the size of the board. Then the image is cropped and scaled to match the MNIST format.

This is the preprocessing flow in order.

```text
PKDrawing
  -> [[CGPoint]]
  -> round-cap grayscale raster
  -> ink bounding-box crop
  -> longest side 20px, aspect ratio 유지
  -> bilinear resample
  -> center of mass를 28×28 중앙으로 이동
  -> [1, 1, 28, 28] Float tensor
```

The Korean in the block says to keep the aspect ratio (유지) and to move the center of mass to the center of the 28×28 image.

The stroke width shown on screen is 6pt for normal input and 3pt for notes. For inference, both are redrawn at 6pt. This keeps the display mode from changing the distribution of the input the model receives.

The preprocessing tests check that empty input returns nil, that the same input gives the same result, that the longest side is 20px, and that the center of mass is in place. They also compare the Core ML label and probability from the Python reference with the Swift result. If any of the crop, scale, pixel direction, or intensity range differs from the training conditions, the model can run without an error and still misrecognize the digit. A successful run does not show that the preprocessing is correct.

## Results to accept and results to drop

The model keeps all 10 classes, 0 through 9. Sudoku has no 0, but I kept it so that an ambiguous circle is not forced into 6 or 9. Input classified as 0 is dropped at the apply step.

The default confidence threshold is 0.8. For notes, where a correction costs less, it is lowered by 0.1 to 0.7. Only when the result is between 1 and 9 and meets the threshold does the app select the cell and place the digit. A 0, a value out of range, low confidence, a model loading failure, a preprocessing failure, and an inference error are all rejects. A reject does not increase the mistake count. It simply does nothing. The ink is cleared regardless of the result.

In the current structure, failures with different causes all merge into a single `.reject`. There is no way to tell whether the model failed to load, the probability was low, or the label was wrong. Counts per cause are needed. I am considering watching the failure rate with on-device statistics only, without sending the original strokes, but that is not implemented yet.

## Results that arrive late

Rendering and prediction run in `Task.detached`, and an accepted result calls `selectCell` and `placeDigit` on the main actor. There are two problems here.

The first is sharing the model. `CoreMLDigitRecognizer` is declared `@unchecked Sendable`, so several tasks can use the same `MLModel` at once. The [`MLModel` documentation](https://developer.apple.com/documentation/coreml/mlmodel) sets a condition that calls on one instance be serialized. `@unchecked Sendable` only skips the compiler check and does not change the runtime condition of `MLModel`. So the current declaration may conflict with that condition.

This is a proposal to let an actor own the recognizer. It is not applied yet.

```swift
actor DigitInference {
    private let recognizer: CoreMLDigitRecognizer

    init() throws {
        recognizer = try CoreMLDigitRecognizer()
    }

    func recognize(_ pixels: [Float]) throws -> DigitRecognition? {
        try recognizer.recognize(pixels: pixels)
    }
}
```

With this, predictions run one at a time, and callers have to use `await`. Cloning the model per queue is another option, but it is something to review only after measurements show that a single actor falls short on latency or throughput.

The second problem is that an old result can be applied after the state has changed. The completion checks only whether the game is in the playing state. After a pause and resume, or after switching note mode, a result requested earlier can still be applied. The two modes can also disagree: the threshold is chosen from the mode at capture time, while `placeDigit` acts on the mode at completion time. I am considering storing the three values below with each request and checking that they still match right before applying.

```text
input revision + target cell + entry mode
```

If the values differ, the result is discarded. A monotonically increasing number works as the revision.

## What automated tests covered and what they did not

The automated tests focus on the pure reducer and renderer. They cover hit testing on the rounded board, the tap and ink conditions, debounce, restart on a distant stroke, normalization, the reject branches, loading the bundled model, the Python and Swift comparison, the allowed touch list, and the delegate wiring.

There are only two handwriting fixtures, 1 and 7. That means only two kinds of input are checked from rendering through to the expected digit. The threshold search in the training script is also a proxy evaluation that uses MNIST with rotation and stroke-width changes, plus 0 and noise. So the recognition rate for real pencil handwriting, and whether 0.8 is the best value, are not proven.

For the delegate, there is a record in a source comment. It diagnosed that on iOS 26.5, PencilKit falls into recursion when the canvas sets itself as its own delegate, so a separate bridge is used. The test only checks whether the canvas and the delegate are the same object. I did not recheck the device logs from that time. Another limit is that the `.pencilOnly` path cannot be exercised with a mouse in the simulator.

Six things need to be checked on devices.

- Precision and reject rate for 1 through 9 across handwriting styles
- Latency of the first response and the order of consecutive results
- The 60pt boundary on each screen size
- Behavior when pencil, finger, pointer, and VoiceOver are mixed
- Discarding earlier results during pause and resume or a mode change
- Whether users can understand a model loading failure or repeated rejects

For metrics, I plan to look at the share of wrong inputs among accepted results, the reject rate, and p95 latency. My judgment was to protect the precision of applied actions first and accept that more inputs go unrecognized.

## What it takes to switch to PKStrokeRecognizer in iOS 27

The iOS 27 beta has [`PKStrokeRecognizer`](https://developer.apple.com/documentation/pencilkit/pkstrokerecognizer). It is an actor, and it is an API that recognizes strokes as text asynchronously on the device. It takes a `PKDrawing` as input. Recognition results can differ by OS, so storing a result means keeping `recognitionVersion` with it. A [guide to recognizing handwriting and converting it to text](https://developer.apple.com/documentation/pencilkit/recognizing-handwriting-and-converting-to-text) is published as well. An app with an iOS 26.5 deployment target needs the Xcode 27 beta SDK, `#available(iOS 27.0, *)`, and a fallback path that uses Core ML.

Swapping only the implementation is hard. The current `DigitRecognizing` takes `[Float]` and returns a digit and a confidence. `PKStrokeRecognizer` takes input through `updateDrawing(_:)`, returns `String?` from `recognizedText(strokeIDs:)`, and has no confidence. So a higher asynchronous boundary that takes the drawing and its context is needed.

```text
recognize(drawing, context) async -> candidate 또는 reject
```

Here 또는 means "or": the call returns either a candidate or a reject.

The Core ML side can keep the current rendering and probability policy, and the new adapter can limit the text to 1 through 9. The rule for applying a result that has no probability automatically needs separate product validation. During the beta period I plan to keep the existing model.

The values 0.8, 0.7, and 60pt in this post are constants set by policy, not optimal values found by measurement. Two handwriting fixtures and a proxy evaluation on MNIST cannot speak for real accuracy, and the six items above stay open until they are checked on devices.
