2026-10-06 • 8 min read • Software Engineer

The OAuth Refresh That Finished Without the App

Our iOS app occasionally logged users out for no visible reason. Nothing in the logs pointed to a bug. The refresh endpoint answered, the tokens were valid, the session was gone anyway.

I found the cause by accident. We had a period of slow backend responses. During one of them I opened the app, switched to another app while it was still loading, and came back later to the login screen. Thinking it through, I realized I had never handled this case: the app gets suspended while a token refresh is in flight.

How Refresh Works

The standard OAuth flow for a mobile app:

  1. The user logs in and the app receives a short-lived access token and a long-lived refresh token.
  2. When the access token expires, the app sends grant_type=refresh_token to the token endpoint.
  3. The server returns a new access token.

Most servers now also return a new refresh token and invalidate the old one. This is refresh token rotation. The OAuth 2.0 Security BCP (RFC 9700) recommends it for public clients. If an old refresh token shows up again, the server treats it as a possible leak and rejects it, often revoking the whole session.

Rotation assumes the client always receives the response. Mobile apps break that assumption.

The Failure

1. Access token expires. App sends refresh with RT1.
2. The network is slow. The user switches to another app.
3. iOS suspends the app a few seconds later.
4. The request still reaches the server. Server consumes RT1, issues RT2.
5. The response arrives at a suspended process. Nobody reads it.
6. The user returns. The request failed with a timeout or connection error.
7. App retries with RT1. Server answers invalid_grant.
8. App clears the session. The user is logged out.

From the server’s point of view, everything worked. From the client’s, the refresh failed. RT2 exists only in the server’s database.

Slow responses made the window large enough to hit by hand. On a good network it is a few hundred milliseconds. It never goes away.

Grace Periods Do Not Close It

Many providers know about this and allow the previous refresh token for a short time:

  • Auth0 has a configurable rotation overlap period.
  • DuckDuckGo’s auth server allows reuse for one hour.
  • Supabase accepts the previous token if it is the parent of the current one, with no time limit.

Time-based grace helps with concurrent refreshes and quick retries. It does not cover a suspended app. The user puts the phone away on a slow train connection and opens the app the next morning. One hour, or three minutes, is long gone.

Supabase’s rule is the robust one: if the client presents the parent of the currently active token, it clearly never received the response, so return the active token again. It does not depend on how long the app was asleep.

You do not control the server’s policy, though, and some servers have no grace at all. The client has to protect itself.

The Fix

Ask iOS for background execution time for the whole refresh, from the request to the Keychain write:

private func refreshWithRetry() async throws -> String {
    let backgroundTask = OAuthRefreshBackgroundTask()
    defer { backgroundTask.end() }

    // send the request, store the rotated tokens, retry on transient errors
}
@MainActor
final class OAuthRefreshBackgroundTask {
    private var identifier: UIBackgroundTaskIdentifier = .invalid

    init() {
        identifier = UIApplication.shared.beginBackgroundTask(
            withName: "OAuth session refresh"
        ) { [weak self] in self?.end() }
    }

    func end() {
        guard identifier != .invalid else { return }
        UIApplication.shared.endBackgroundTask(identifier)
        identifier = .invalid
    }
}

The same sequence with the fix:

1. Access token expires. App begins a background task, sends refresh with RT1.
2. The network is slow. The user switches to another app.
3. iOS keeps the app running because of the background task.
4. The request reaches the server. Server consumes RT1, issues RT2.
5. The app reads the response and stores RT2.
6. The app ends the background task. iOS suspends it.
7. The user returns hours later. App refreshes with RT2.
8. Server issues RT3. The user stays logged in.

Steps 1, 2 and 4 are unchanged. The difference is step 3: the app is still running when the response arrives, so RT2 does not exist only in the server’s database.

Two related changes:

  • Stop retrying on invalid_grant. Retrying a consumed token never succeeds.
  • Do not clear the session on transport errors. A timeout says nothing about the token.

What It Covers

iOS gives a background task roughly 30 seconds, counted from the moment the app leaves the foreground.

The refresh starts. Four seconds later the user switches to another app. The server answers at second fifteen. The app is still running, reads the response, and stores RT2. Then it is suspended. The user comes back two hours later, and the next refresh uses RT2. RT2 has never been used, so the grace period does not matter.

If the response is lost inside the window, for example because the connection dropped, the retry with RT1 happens within seconds. Any grace period of a minute or more accepts it.

What It Does Not Cover

The response arrives after the window. The app is suspended before reading it. The retry waits until the user comes back, which is usually long after any time-based grace period. Two options remain:

  • Parent-token reuse on the server, as Supabase does it.
  • Sending the refresh through a background URLSession. The system completes the transfer while the app is suspended and delivers the response later. The cost: upload tasks from a file, delegate callbacks instead of async/await, and cancellation if the user force-quits the app.

The server has no grace at all. A lost response cannot be retried. Only reading the first response saves the session.

The device is locked. Users often leave the app and lock the phone. About ten seconds after the lock, Keychain items stored with kSecAttrAccessibleWhenUnlocked become unavailable. The response arrives, the write of RT2 fails, and the token is lost. Our session item used exactly that class. Tokens that must be written in the background need kSecAttrAccessibleAfterFirstUnlockThisDeviceOnly. A locked Keychain is not a reason to clear the session either. It says nothing about the token.

App extensions. They cannot use UIApplication. Use ProcessInfo.performExpiringActivity there.

Other Apps

After fixing it, I checked open-source iOS apps and SDKs for the same pattern: rotation on the server, no background execution around the refresh on the client. I read the code; I did not reproduce the bug in these apps.

This is not a list of shame. The problem is hard to see intuitively. Once you know about it, it looks logical and almost obvious. I had not handled it in any of the apps I built before this one, and there were several.

DuckDuckGo. The subscription token refresh runs in a plain Task on an ephemeral URLSession, with no background task in the client. The server allows reuse for one hour. The app does start background tasks when it leaves the foreground, but none of them covers the refresh. A 15-second “App Background Safety Net” sits behind the genericBackgroundTask remote flag, which is disabled in the production config. On iOS 26, a Screen Time cleanup task starts on every transition, but it ends as soon as the cleanup returns, usually within milliseconds. After that, the app is suspended as usual. On a rejected token, the app tries to restore an App Store purchase. Otherwise it signs the user out.

Proton. The shared protoncore_ios library refreshes through /auth/v4/refresh, gets a new refresh token, and logs the user out on a 400 or 422 response. Its only background task helper is used in signup. Proton Pass and Proton VPN, both built on this library, have no background task anywhere in their code. I could not verify Proton’s server reuse policy. The new Proton Mail app uses a Rust core and starts a background task on every transition to background, so it is likely covered.

Auth0.swift. CredentialsManager renews credentials on URLSession.shared with no lifecycle handling. With rotation enabled, a token reused outside the overlap period is rejected. The SDK’s own docs recommend an overlap of at least 180 seconds. Any app using it inherits the problem unless it wraps the call itself.

OAuthenticator (ChimeHQ). No background task, and any refresh error clears the login, including a timeout. It supports Bluesky, whose OAuth server deletes the session when an old refresh token is replayed.

Some get it right. Okta’s SDK wraps refresh in a background task. Infomaniak kDrive has a comment saying the app must stay awake while refreshing.

The fix is about ten lines. The bug is rare, invisible in logs, and looks like a random logout to the user. That combination is why it survives.

If your app has a small group of users with random auth problems that nobody has managed to trace, this might be the case.