Task — engineering-spec@1

"SSO JIT provisioning must take the person's name from the ID token, not their email address"

doneTASK-AUTH-111
module auth · class product · priority p1 · created 2026-07-11 · shipped 2026-07-12
depends on none · blocks none

§1 - Description (BCP-14 normative)

  1. The OIDC callback MUST parse the person's name from the verified ID token, using the standard OpenID Connect claims: name, and failing that given_name + family_name, and failing that preferred_username.
  1. At JIT provisioning the service MUST bind subjects.display_name from that resolution chain, falling back in order to: the ID token's name; given_name family_name joined; preferred_username; the local-part of the email; and only then the full email address. It MUST NOT bind display_name directly from the email, which is what it does today (oidc.rs binds idp_email.unwrap_or("") into the display_name column - the bug this task exists to remove).
  1. The service MUST NOT invent, prettify, title-case, or otherwise transform a name it was given. If the IdP says the person is called nguyenvana, that is their name. Guessing at capitalisation or splitting on separators is how a product mangles Vietnamese names.
  1. On the ON CONFLICT (tenant_id, handle) DO UPDATE path - a returning SSO user - the service MUST refresh display_name only when the stored value was never really a name: that is, when it is NULL, empty, or byte-equal to the subject's email. It MUST NOT overwrite a display_name that differs from the email, because that value was set deliberately - by an administrator, by a future profile editor, or by a data repair - and silently reverting it on next login is worse than the original bug.
  1. Consequently the fix MUST be self-healing: every existing subject whose display_name is currently their email address acquires their real name the next time they sign in, with no migration and no administrator action.
  1. The same defect exists in saml.rs, which binds display_name the same way. It MUST be fixed in the same change, with the same resolution chain and the same no-clobber rule. Fixing one and not the other leaves a second door open and guarantees the bug is rediscovered.
  1. The migration MUST NOT attempt to reconstruct names for existing rows. There is no source of truth for them in our database - the ID token was discarded - and inventing a name from an email local-part would produce van-anh.vu where the person is called Vũ Vân Anh. §1 #5's self-healing path is the correct repair.
  1. The resolution MUST be logged at debug with the rung of the chain that matched, never with the name itself. tracing::debug!(rung = "given_name+family_name") is a diagnostic; tracing::debug!(?name) puts personal data in the log stream, which the privacy policy says we do not do.
  1. No claim beyond those in §1 #1 MAY be persisted. In particular the picture claim MUST NOT be stored: the published privacy policy enumerates what we collect, and adding a field to subjects is a change to that enumeration, not an implementation detail. If we want avatars, that is a separate task and a privacy-policy revision, in that order.

§2 - Why this design (rationale for humans)

Why this is a defect and not a cosmetic issue. oidc.rs line ~965 binds the email into the display_name column. Every human provisioned by Google SSO therefore renders as van-anh.vu@cyberskill.world wherever a name belongs: in the channel list, above every message they have ever sent, in mentions, in the member picker, and in any screenshot or export. It reads as broken software. It also means the address is displayed in contexts where only a name was intended - a small but real leak, and one that has already forced a manual UPDATE against production to make a store screenshot presentable.

Why the code contradicts the published policy. cyberskill.world/en/cyberos/privacy says, in terms: "When you sign in with Google we receive your name, your email address, and your Google account identifier." We do receive the name. We then drop it on the floor and store the email in its place. The policy is not wrong about what Google sends us; the code is wrong about what we do with it. Aligning them is the whole of this task.

Why the no-clobber rule (§1 #4) is a MUST and not a nicety. The naive fix - always refresh display_name from the ID token on login - would work today and break the moment anything else can set a name. An administrator who fixes a colleague's name by hand, or a future profile editor, would watch their change silently revert on that person's next sign-in, with no error and no trace. Refreshing only the sentinel value (null, empty, or exactly the email) is the narrow rule that repairs the damage without acquiring the authority to undo a human's decision.

Why not backfill (§1 #7). Because the information is gone. We can see that display_name = email, so we can see which rows are wrong, but we cannot see what they should say - the ID token that carried the name was never persisted. The only honest sources are the IdP (which means a login) or a human (which means typing). Deriving Van Anh Vu from van-anh.vu@ looks like a fix and is a guess, and it guesses wrong on exactly the names that matter most to a Vietnamese company. Self-healing on next login is slower and correct.

Why picture is explicitly excluded (§1 #9). It is right there in the ID token, it would make the product nicer, and adding it would quietly make the published Data Safety declaration false. The store form and the policy both enumerate what we collect. Widening that set is a decision with paperwork attached, not a line of code.

§3 - API contract

The claims we read

// services/auth/src/oidc.rs

/// Standard OIDC claims. We read only what §1 #1 names. `picture` is deliberately absent:
/// adding it here would silently widen the Data Safety declaration (§1 #9).
#[derive(Debug, Deserialize)]
struct IdTokenProfile {
    #[serde(default)] name: Option<String>,
    #[serde(default)] given_name: Option<String>,
    #[serde(default)] family_name: Option<String>,
    #[serde(default)] preferred_username: Option<String>,
}

The resolution chain

/// §1 #2. Ordered, total, and deliberately dumb: no transformation of any value it is handed (§1 #3).
/// Returns the rung that matched alongside the value, so the caller can log the rung without logging
/// the name (§1 #8).
fn resolve_display_name<'a>(
    p: &'a IdTokenProfile,
    email: Option<&'a str>,
) -> (&'static str, String) {
    if let Some(n) = p.name.as_deref().map(str::trim).filter(|s| !s.is_empty()) {
        return ("name", n.to_string());
    }
    let given  = p.given_name.as_deref().map(str::trim).filter(|s| !s.is_empty());
    let family = p.family_name.as_deref().map(str::trim).filter(|s| !s.is_empty());
    match (given, family) {
        (Some(g), Some(f)) => return ("given_name+family_name", format!("{g} {f}")),
        (Some(g), None)    => return ("given_name", g.to_string()),
        (None, Some(f))    => return ("family_name", f.to_string()),
        (None, None)       => {}
    }
    if let Some(u) = p.preferred_username.as_deref().map(str::trim).filter(|s| !s.is_empty()) {
        return ("preferred_username", u.to_string());
    }
    match email {
        Some(e) if e.contains('@') => ("email_local_part", e.split('@').next().unwrap().to_string()),
        Some(e)                    => ("email", e.to_string()),
        None                       => ("handle", String::new()),   // caller substitutes the handle
    }
}

The JIT insert

let (rung, display_name) = resolve_display_name(&profile, idp_email);
tracing::debug!(target: "cyberos_auth::oidc", rung, "resolved display_name");   // §1 #8 - rung, not name

let row: (Uuid,) = sqlx::query_as(
    "INSERT INTO subjects (tenant_id, handle, display_name, email, kind, status, password_hash, roles)
          VALUES ($1, $2, $3, $4, 'human', 'active', $5, $6)
     ON CONFLICT (tenant_id, handle) DO UPDATE
        SET email = COALESCE(EXCLUDED.email, subjects.email),
            -- §1 #4: refresh ONLY a display_name that was never really set. A value that differs
            -- from the email was put there on purpose; do not silently revert someone's decision.
            display_name = CASE
                WHEN subjects.display_name IS NULL
                  OR subjects.display_name = ''
                  OR subjects.display_name = subjects.email
                THEN EXCLUDED.display_name
                ELSE subjects.display_name
            END,
            updated_at = NOW()
   RETURNING id",
)
.bind(idp.tenant_id)
.bind(&handle)
.bind(&display_name)          // <- was: idp_email.unwrap_or("")   THE BUG
.bind(idp_email)
.bind(&sso_password_hash)
.bind(&idp.default_roles)
.fetch_one(&mut *tx)
.await
.map_err(internal)?;

Migration

-- services/auth/migrations/0026_sso_display_name_backfill.sql
-- TASK-AUTH-111. There is deliberately NO backfill of display_name here (§1 #7): the ID token that
-- carried the person's name was never persisted, so we can see WHICH rows are wrong but not what
-- they should say. Deriving "Van Anh Vu" from "van-anh.vu@" is a guess, and it guesses wrong on
-- exactly the names that matter most to a Vietnamese company. Every affected subject self-heals on
-- their next sign-in via the ON CONFLICT rule.
--
-- What this migration DOES is make the damage visible and measurable, so we can watch it drain.

CREATE OR REPLACE VIEW subjects_display_name_unset AS
    SELECT id, tenant_id, handle, email, created_at
      FROM subjects
     WHERE kind = 'human'
       AND email IS NOT NULL
       AND (display_name IS NULL OR display_name = '' OR display_name = email);

COMMENT ON VIEW subjects_display_name_unset IS
    'TASK-AUTH-111: humans still wearing their email as a display name. Drains to zero as people sign in.';

GRANT SELECT ON subjects_display_name_unset TO cyberos_ro;

§4 - Acceptance criteria

  1. name wins - an ID token carrying name: "Trịnh Thái Anh" provisions a subject whose display_name is exactly Trịnh Thái Anh, diacritics intact, untransformed.
  2. given_name + family_name is the second rung - a token with no name but both parts provisions "<given> <family>".
  3. A single part is used alone - a token with only given_name provisions that value, not an empty-padded join.
  4. preferred_username is the third rung - used when no name claims are present.
  5. The email local-part is the fourth rung - a token with no name claims at all and email: a.b@c.com provisions a.b, never a.b@c.com.
  6. display_name is never the full email - across every rung, no provisioned subject has display_name = email.
  7. A returning user with an email as their name is repaired - a subject stored with display_name = email acquires their real name on the next sign-in, with no migration.
  8. A deliberately set name is never clobbered - a subject whose display_name was set to "Play Review" by hand keeps it across sign-ins, even though the ID token would resolve to something else.
  9. A blank ID-token name does not blank the stored name - a token whose name is "" or whitespace falls through the chain and never writes an empty display_name.
  10. The name is never logged - no log line at any level contains the resolved name; the debug line contains only the rung that matched.
  11. picture is not persisted - no column receives it, and the subjects row after a Google sign-in has no avatar data.
  12. SAML is fixed identically - the same eight assertions hold against saml.rs's provisioning path.
  13. The visibility view drains - subjects_display_name_unset returns a row for an affected subject before their next sign-in and no row after it.

§5 - Verification

// services/auth/tests/oidc_jit.rs

#[tokio::test]
async fn the_resolution_chain_walks_in_order() {                     // AC 1-5, 9
    assert_eq!(resolve(tok().name("Trịnh Thái Anh")).1, "Trịnh Thái Anh");
    assert_eq!(resolve(tok().given("Thái Anh").family("Trịnh")).1, "Thái Anh Trịnh");
    assert_eq!(resolve(tok().given("Thái Anh")).1, "Thái Anh");
    assert_eq!(resolve(tok().preferred("stephen")).1, "stephen");
    assert_eq!(resolve(tok().email("van-anh.vu@cyberskill.world")).1, "van-anh.vu");
    // whitespace-only claims fall through rather than writing a blank
    assert_eq!(resolve(tok().name("   ").email("a.b@c.com")).1, "a.b");
}

#[tokio::test]
async fn display_name_is_never_the_full_email() {                    // AC 6 - the bug, pinned
    let app = harness().await;
    for tok in [tok().name("A B"), tok().preferred("ab"), tok().email("a.b@c.com")] {
        let s = app.sso_login(tok.email("a.b@c.com")).await;
        assert_ne!(s.display_name, s.email.unwrap(),
                   "display_name must never be bound from the email column");
    }
}

#[tokio::test]
async fn a_returning_user_self_heals_but_a_set_name_survives() {     // AC 7, 8 - the no-clobber rule
    let app = harness().await;

    // Damaged row, exactly as production has them today.
    let s = app.seed_subject(display_name = "van-anh.vu@cyberskill.world",
                             email        = "van-anh.vu@cyberskill.world").await;
    app.sso_login(tok().sub(s.idp_sub).name("Vũ Vân Anh")).await;
    assert_eq!(app.subject(s.id).await.display_name, "Vũ Vân Anh");   // repaired

    // Deliberately set row: an admin fixed this by hand. It must survive.
    let p = app.seed_subject(display_name = "Play Review",
                             email        = "play-review@cyberskill.world").await;
    app.sso_login(tok().sub(p.idp_sub).name("play-review")).await;
    assert_eq!(app.subject(p.id).await.display_name, "Play Review");  // NOT clobbered
}

#[tokio::test]
async fn the_name_never_reaches_the_log_stream() {                   // AC 10
    let logs = capture_logs(Level::DEBUG, || async {
        harness().await.sso_login(tok().name("Trịnh Thái Anh")).await;
    }).await;
    assert!(logs.contains(r#"rung="name""#));
    assert!(!logs.contains("Trịnh Thái Anh"), "personal data must not reach the log stream");
}

#[tokio::test]
async fn picture_is_not_persisted() {                                // AC 11
    let app = harness().await;
    let s = app.sso_login(tok().name("A B").picture("https://…/photo.jpg")).await;
    let raw = app.raw_subject_row(s.id).await;
    assert!(!raw.to_string().contains("photo.jpg"));
    assert!(!raw.to_string().contains("picture"));
}

#[tokio::test]
async fn saml_is_fixed_too() {                                       // AC 12
    let app = harness().await;
    let s = app.saml_login(assertion().display_name("Trịnh Thái Anh")
                                      .email("thai-anh.trinh@cyberskill.world")).await;
    assert_eq!(s.display_name, "Trịnh Thái Anh");
}

#[tokio::test]
async fn the_visibility_view_drains() {                              // AC 13
    let app = harness().await;
    let s = app.seed_subject(display_name = "a.b@c.com", email = "a.b@c.com").await;
    assert_eq!(app.view_rows("subjects_display_name_unset").await.len(), 1);
    app.sso_login(tok().sub(s.idp_sub).name("A B")).await;
    assert_eq!(app.view_rows("subjects_display_name_unset").await.len(), 0);
}

§6 - Implementation skeleton

(§3 is the skeleton. The change is three lines of binding, one CASE in the upsert, one helper, and the same again in saml.rs.)

§7 - Dependencies

§8 - Example payloads

The Google ID token, as verified today:

{
  "iss": "https://accounts.google.com",
  "sub": "104…",
  "email": "thai-anh.trinh@cyberskill.world",
  "email_verified": true,
  "name": "Trịnh Thái Anh",
  "given_name": "Thái Anh",
  "family_name": "Trịnh",
  "picture": "https://lh3.googleusercontent.com/…"
}

The subject row today (the bug):

{ "handle": "@thai-anh.trinh",
  "display_name": "thai-anh.trinh@cyberskill.world",
  "email": "thai-anh.trinh@cyberskill.world" }

The subject row after this task:

{ "handle": "@thai-anh.trinh",
  "display_name": "Trịnh Thái Anh",
  "email": "thai-anh.trinh@cyberskill.world" }

§9 - Open questions

Deferred:

§10 - Failure modes inventory

FailureDetectionOutcomeRecovery
display_name bound from the email columnAC 6 asserts display_name != email on every rungThe defect in production todayThe bind is changed; the AC fails if anyone reverts it
SSO overwrites a hand-set name on next loginAC 8 seeds a deliberately set name and signs inAn admin's repair silently reverts, with no errorCASE refreshes only null / empty / equal-to-email
A blank name claim writes an empty display nameAC 9; every rung trims and filters emptyA person renders as nothing at allThe chain falls through on whitespace
Backfill invents a wrong name§1 #7; no backfill existsVan Anh Vu where the person is Vũ Vân AnhThe migration creates a view, not an UPDATE
Name written to the log streamAC 10 greps captured logs for the namePersonal data in logs, contradicting the privacy policyThe debug line carries the rung only
picture quietly persistedAC 11 greps the raw rowThe published Data Safety declaration becomes falseIdTokenProfile has no picture field to deserialise into
SAML left brokenAC 12 runs the same assertions against saml.rsThe bug is rediscovered through the other doorBoth call sites are changed in one commit
Diacritics mangled by a "helpful" transformationAC 1 asserts byte equality on Trịnh Thái AnhVietnamese names corrupted - the worst possible failure for this company§1 #3 forbids transformation; the resolver copies
A future contributor "simplifies" the CASE to an unconditional SETAC 8 failsThe no-clobber guarantee is lostPinned by an assertion, not a comment
Subject has no email and no name claimsChain's final rung returns empty; caller substitutes the handleThe person renders as their handle, not as blankExplicit final arm in resolve_display_name

§11 - Implementation notes

End of TASK-AUTH-111.

Audit

§1 - Verdict summary

TASK-AUTH-111 removes a defect found live: oidc.rs binds the subject's email into the display_name column at JIT provisioning, so every Google-SSO user wears their email address as their name. 9 normative clauses, one migration (a visibility view, deliberately no backfill), 13 acceptance criteria, 7 tests, a 10-row failure-mode inventory. Six findings raised; all resolved. This is a small surface with two sharp edges: the no-clobber rule (§1 #4) and the refusal to guess a name (§1 #7).

§2 - Findings (all resolved)

ISS-001 - The naive fix would silently revert an administrator's repair

The obvious implementation refreshes display_name from the ID token on every login. It works today, and it breaks the moment anything else can set a name - which is already the case: a manual UPDATE was run against production during the Play submission work to fix exactly one of these rows. Under the naive fix, that repair would silently revert on the next sign-in, with no error and no trace. Resolved: §1 #4 refreshes only a display_name that is NULL, empty, or byte-equal to the email - the sentinel value the bug itself writes, and one no human would choose. §3 implements it as a CASE inside the ON CONFLICT DO UPDATE. AC #8 seeds a deliberately-set name ("Play Review") and asserts it survives a login whose token would resolve to something else. §10 rows 2 and 9.

ISS-002 - The migration wanted to backfill, and would have guessed wrong

The first draft derived a name from the email local-part for existing rows: van-anh.vu@ becomes Van Anh Vu. It looks like a repair and it is a guess, and it guesses wrong on precisely the names that matter most to a Vietnamese company - the person is Vũ Vân Anh. Worse, it would overwrite the sentinel, so the self-healing path in §1 #4 would never fire and the wrong name would become permanent. Resolved: §1 #7 forbids reconstructing names; §2 explains why; the migration ships a view (subjects_display_name_unset) instead, so the damage is countable and can be watched draining to zero as people sign in. AC #13. §10 row 4.

ISS-003 - SAML has the same bug and was going to be left open

saml.rs binds display_name the same way as oidc.rs. Fixing one door and leaving the other means the bug is rediscovered by whoever next onboards a SAML tenant, having read a changelog entry that says it was fixed. Resolved: §1 #6 makes both call sites normative in this task; AC #12 runs the same assertions against the SAML path.

ISS-004 - The debug line would have logged the person's name

Adding tracing::debug!(?display_name) is the reflex when debugging a name-resolution bug, and it puts personal data into the log stream - contradicting the privacy policy this task exists to align the code with. Removing the log line entirely is the other reflex, and leaves the resolver unobservable. Resolved: §1 #8 logs the rung of the chain that matched, never the value; §3's resolve_display_name returns (&'static str, String) so the caller gets both properties for free; AC #10 captures the log stream and asserts the name is absent while the rung is present. §11.

ISS-005 - picture was quietly in scope

The picture claim sits in the same ID token, avatars would make the product nicer, and adding the column is one line. It would also make the Data Safety declaration we are about to file with Google false, since that form and the published privacy policy both enumerate what we collect. Resolved: §1 #9 forbids persisting any claim beyond those in §1 #1, names picture explicitly, and states that widening the set is a policy revision first and a code change second. IdTokenProfile has no picture field to deserialise into. AC #11 greps the raw row. §9 keeps it as a deferred task with its paperwork attached.

ISS-006 - Whitespace-only and empty claims would have written a blank name

name: "" and name: " " are both things IdPs emit. The draft's Option check treated them as present, so the person would have rendered as nothing at all - a worse outcome than the email. Resolved: every rung of §3's chain applies trim() and filter(|s| !s.is_empty()) before accepting, so blanks fall through; AC #9 asserts a whitespace name falls through to the email local-part rather than writing an empty display_name. §10 row 3.

§3 - Resolution

Six findings, all resolved. The two that would have caused lasting damage - ISS-001 (reverting a human's repair) and ISS-002 (permanently inventing a wrong Vietnamese name) - are each pinned by an acceptance criterion that fails loudly, not by a comment. ISS-005 is the one worth re-reading before anyone "just adds avatars": the code change is trivial and the paperwork is not.

Score = 10/10.


End of TASK-AUTH-111 audit.

Ship record (2026-07-12 - status-drift reconciliation)

  • Implemented at bc7af7b (parallel session); 9/9 clause verification PASS (packet: docs/tasks/.workflow/TASK-AUTH-111/review-packet.md).
  • Test evidence: 6 unit tests incl. anti-prettify canary + picture-cannot-land proof; operator confirmed tests green (CI/cargo) - sandbox has no Rust toolchain, gap named.
  • HITL: operator verdict 2026-07-12 in-chat "Tests green - approve + done" (both gates).