Eunoia Lab · Working Paper · Study 1 Pilot

A pilot of the systematic walkthrough

One reader, four real interfaces, the four commitments, scored twice: once leniently, once strictly. Where the two readings part is where the instrument is loose.

Abstract

The Confidant Standard proposes Study 1 but does not run it: two or more trained raters, a shared codebook, a corpus of live interfaces, a reliability figure at the end. Before committing anyone to that, I wanted to know whether the codebook holds up in the hand. So I ran a smaller thing first, an audit of the instrument by a single analyst. I took four real interfaces, one per archetype, and scored each against the four commitments twice over: once giving every signal the benefit of the doubt, once withholding credit until the signal was clearly doing work. The point was never to measure the market, and this is not a reliability test. It is a way to find where the codebook is loose, and it is loose in a useful way. The lenient and strict readings settled on the same call for Restraint almost everywhere and parted repeatedly on the other three, always at the seam between partial and absent. Each place they part names something the codebook needs to fix before a real rater is ever recruited.

How to read this

Study 1 proper needs independent human raters, and this pilot does not have them. It has one reader, me, scoring each interface twice: a lenient pass that gives a signal the benefit of the doubt, and a strict pass that withholds credit until the signal is clearly load-bearing and defaults otherwise to the incentives. Setting those two readings against each other flushes out the borderline calls, which is the whole point, but it is not inter-rater agreement and I report no agreement figure. This is a structured self-critique of the instrument, not a measurement of it.

1The corpus

One interface per archetype in Table 2 of the paper. I checked the first three against how they actually behave today, and read the fourth at the level of its mechanics, since there is no single feature to point at.

2The codebook, in brief

The scale is the one from Section 5. Present when the signal is clearly there and carrying weight, partial when it exists but is weak, incidental, or inconsistent, absent when it is not there at all. The observable anchors are unchanged from Table 1.

3The two readings

Both readings sit in each cell, the lenient one first, the strict one second. Where they part, the cell carries a gold rule and a warm tint.

Table P1. Lenient and strict readings of four interfaces
Interface Restraint Legibility Latitude Non-exploitation
Sephora VA / Color IQ Archetype A lenstr lenstr lenstr lenstr
Beauty Genius Archetype B lenstr lenstr lenstr lenstr
Skin Genius Archetype C lenstr lenstr lenstr lenstr
Pinterest feed Archetype D lenstr lenstr lenstr lenstr
present partial absent readings part

4Where the two readings part

The pattern worth noting is not how often the two readings matched, which says more about how I set them up than about the interfaces. It is the shape of where they came apart. The lenient and strict passes never landed on opposite ends of the scale. Nowhere did one call a signal present and the other absent. Every place they parted was a single step, present against partial or partial against absent, and nearly all of it gathered at that lower seam. The two readings are not disagreeing about whether a signal is there. They are disagreeing about how much of it has to be there before it counts, which is a question about the codebook, not about the interfaces.

It does not spread evenly across the four commitments either. Restraint held steady: the two readings matched on it almost everywhere, nearly always on a flat absent, which tells me the construct is doing clean work and I can more or less trust it as written. The instability pooled in the other three. Legibility, Latitude, and Non-exploitation parted under nearly every interface, and they parted the same way each time. The lenient reading saw a capability and gave partial; the strict reading saw no evidence it was load-bearing and gave absent. That is not a disagreement about what the interface does. It is an unsettled definition of what partial means, and the four repairs below are how I would settle it before anyone else scores a thing.

5What the disagreements diagnose

Each split points somewhere specific. Four repairs, roughly in the order I would make them.

  1. The hole where there is no affect at all. Sephora's try-on cannot exploit a low moment because it never registers one. The generous reader called that partial and meant low risk; the strict reader called it absent and meant there is no guardrail here, and a missing channel is not a virtue. Both are right, which is the tell that the scale is short a code. Non-exploitation should either get an explicit not-applicable, or a rule that it is only creditable where affect is actually being read. Without it the column quietly files safe-by-design and safe-by-accident under the same mark, and that is the one distinction the whole commitment exists to make.
  2. The gap between showing what and showing why. Skin Genius and Beauty Genius both put their detected attributes on screen, and to a generous eye that reads as legibility. But the anchor asks for reasoning a person can push back on, and a score you can only nod at is not that. The observable needs splitting: disclosing what the system found is the weak version, and giving a reason you could argue with is the real one.
  3. Whether running the model again counts as latitude. A retake every six weeks, a tap that says show me fewer like this, both feed fresh input to the same model of you. The generous reader took these as room to move; the strict reader said they refresh the inputs without ever letting you contradict the read. I side with the strict reading, and the codebook should say so. Latitude means overriding the model's picture of you and watching that stick, not re-running the same inference on a new photo.
  4. Capability against behavior. A chatbot can decline in principle. It rarely does when every path is pointed at a purchase. Credit the no you can observe, never the one the system is theoretically capable of.
If one thing here survives a proper run, it will be this: Restraint is the cleanest of the four to score and Non-exploitation the muddiest, and the muddiness is mostly an artifact of the codebook, the kind you can engineer out before a single rater is recruited.

6Limitations

One analyst, not two, so both readings are the same judgment in two moods and none of this is reliability, nor is it offered as such. Four interfaces, one look each. And this was reading, not use. I never sat in front of a system while genuinely feeling bad about my own face, so every Non-exploitation mark is inferred from how the thing is built rather than watched in the moment it would matter. The feed I only ever read at the level of mechanics.

A real Study 1 fixes all of that: trained human raters, more than one of them, scoring independently so an agreement figure actually means something; the codebook with these four repairs folded in; several instances per archetype instead of one; and, for Non-exploitation at least, a scripted low-confidence moment so the commitment is observed under load rather than guessed at from the schematics. What this pilot gives me is not a finding. It is a better instrument to hand the people who will run the thing for real.