EVA · Field Testing

Tester Issue Log

Tester report & 9 screenshots · sessions 19:05–19:13 and 15:22–15:23 · Correspondent, Memory Search, Scenario Analysis, Cue Cards, Cartographer chat
Account context: secondary email connected · principal domain axiometis.com · primary inbox indexing status unconfirmed (see M-3)
4 BLOCKER  ·  9 HIGH  ·  4 MEDIUM  ·  2 LOW  ·  1 OPEN QUESTION  ·  2 NOT DEFECTS
■ Most damaging · P-1 · EVA does not know who it works for
The principal's connected Gmail is on axiometis.com, their company is Axiometis — and EVA asked them to explain their relationship to their own employer.

Asked “how does it impact axiometis” in The Cartographer chat, EVA replied:

“I don't have any information about Axiometis in memory yet. This appears to be a new entity that hasn't been tracked.”

… then asked the principal: “1. What is Axiometis? (company, competitor, partner, client?) 2. What's your relationship to them? 3. What space do they operate in?”

This is the anchor the rest of the product scores against. Without it, signal ranking, scenario analysis, drift detection and cue-card selection have nothing to be relevant to — which is why the same session shows Signals 198 alongside an assistant that cannot relate any of them to the principal's business. It also directly contradicts the “sworn to one principal” premise the product is built on.

Corroboration: Scenario Analysis earlier accepted “What should AxioMetis work on…” and emitted confident probability bands without ever knowing what AxioMetis was. The chat surface, asked about the same entity, correctly reports it has never heard of it. The scenario engine was not analysing — it was emitting defaults.

EVA · Issue Index

Summary

Every entry below is grounded in an exact on-screen observation from the tester's screenshots, quoted verbatim. Where a likely cause is offered it is labelled as such and flagged for verification — none of the root-cause notes have been confirmed against a running system.

IDAreaSeverityIssue
C-1CorrespondentHighAll queue cards render identically — no sender, subject, or timestamp
C-2CorrespondentLowInternal enum DRAFT_AND_HOLD shown as user-facing text
C-3CorrespondentBlockerCard says “draft prepared, awaiting approval” but opens an empty dictation screen; draft never shown
C-4CorrespondentHighRaw Gmail thread ID is the only identifier on the detail screen
C-5CorrespondentHighOriginal message not shown — user replies blind
C-6CorrespondentHighNo approve / reject / edit controls anywhere; only “record”
M-1Memory SearchBlockerReturns NO RESULTS for a query with a known correct answer
M-2Memory SearchMediumNatural-language question accepted, but empty state gives no reason or fallback
M-3Memory SearchQuestionTarget email may sit in the primary inbox, which may not be connected
S-1Scenario AnalysisBlockerProbability bands are static boilerplate, unrelated to input
S-2Scenario AnalysisHigh“Key signals” are unrelated crawled news headlines
S-3Scenario AnalysisMediumSupplied URL in the situation text is ignored
S-4Scenario AnalysisHighRisk classifier reports “no risk markers” for an explicitly war-risk scenario
Q-1Cue CardsMediumRaw internal UUIDs rendered to the user
Q-2Cue CardsHighBoilerplate scenario output promoted into the brief as a PATTERN signal
Q-3Cue CardsLowEvery card shows MEDIUM priority — no discrimination
P-1Principal contextBlockerEVA does not know the principal's own employer, and asks them to explain it
P-2Principal contextHighNothing derives organisation from the connected email domain
P-3CartographerHigh198 signals ingested, none mappable to the principal's business
P-4CartographerMediumScenarios 0 despite a scenario having been run and surfaced as a cue card
E-1AuthExpectedGoogle marks the app untrusted (unverified OAuth)

Cross-Cutting Read

Two root gaps account for most of the list. The second is the more damaging in a product whose entire positioning is trustworthiness.

1
No principal context — nothing to be relevant to.

EVA does not know who the principal is or who they work for (P-1, P-2). Everything requiring relevance — signal ranking, scenario analysis, drift, cue-card selection — has nothing to score against, so it falls back to generic output (S-2, P-3) or fixed defaults (S-1, S-4). Fixing signal relevance without fixing this is not possible; the ranking has no anchor.

2
States asserted without the substance behind them.

The Correspondent queue claims drafts that don't render (C-3). The scenario engine emits confident probability bands for an entity it has never heard of (S-1, corroborated by P-1). The risk classifier reports “no risk markers” on an explicitly war-risk premise (S-4). Boilerplate is promoted into the brief as a finding (Q-2). In every case the UI communicates more confidence than the underlying system has earned.

Correspondent

C-3

“Draft prepared” opens an empty dictation screen

Blocker
ObservedCard states “draft prepared, awaiting principal approval.” Tapping it opens a screen titled DICTATE REPLY containing only:
Thread: 19fd2385b588dff8 Speak your reply — transcribed silently (no auto-send). [mic] Press and hold to record
The prepared draft is not displayed anywhere.
ExpectedThe draft text, with approve / reject / edit actions.
ImpactThe core Correspondent loop is unusable. The list asserts work was done; the detail screen asks the user to do that work themselves.
Open questionAre drafts being generated and not rendered, or is the “draft prepared” state set without a draft ever being created? These need different fixes and the screenshots cannot distinguish them.
Evidence — screenshots 1, 2, 3
C-1

Queue cards are undifferentiated

High
ObservedUnder NEEDS ATTENTION, seven consecutive cards render byte-identical:
DRAFT_AND_HOLD action required: draft prepared, awaiting principal approval
No sender, subject, snippet, timestamp, or priority on any card.
ExpectedEach item identifiable at a glance — counterparty, subject or summary, age.
ImpactTriage is impossible; the user cannot choose which item to open or tell whether two cards are the same thread.
Evidence — screenshot 1
C-4

Raw thread ID as the only identifier

High
ObservedThread: 19fd2385b588dff8 and Thread: 19fcbbc5037d78d1 — raw Gmail thread identifiers presented as the screen's title content.
ExpectedSubject line and counterparty.
Evidence — screenshots 2, 3
C-5

No view of the message being replied to

High
ObservedThe dictation screen shows no original message, sender, or thread history.
ExpectedAt minimum the message being replied to.
ImpactEven treating dictation as the intended flow, the user is composing blind.
Evidence — screenshots 2, 3
C-6

No approval controls exist

High
ObservedThe only interactive element is the record button. No approve, reject, edit, or send control is present on either the card or the detail screen.
ExpectedGiven the stated state (“awaiting principal approval”), an approve/reject path.
NoteThe copy says “no auto-send”, but there is no visible send path at all, so the post-dictation flow is unclear.
Evidence — screenshots 1, 2, 3
C-2

Internal enum shown as UI copy

Low
ObservedDRAFT_AND_HOLD displayed as the card's headline label.
ExpectedHuman-readable state (“Draft ready for approval”).
Evidence — screenshot 1

Memory Search

M-1

No results for a query with a known answer

Blocker
ObservedQuery: “when did I last communicate with Indhu”NO RESULTS. The tester states the correct answer is roughly three days ago, and that EVA could not locate the relevant email.
ExpectedThe matching communication, with a date.
ImpactFeature returns nothing for its primary use case.
Evidence — screenshot 4
M-2

Empty state gives no reason and no fallback

Medium
ObservedA bare NO RESULTS.
ExpectedDistinguish nothing indexed from nothing matched; offer partial matches, a scope hint, or a way to widen the search.
NoteThe input invites a natural-language question while the backend is fact/keyword retrieval. If those don't match, either the prompt or the retrieval needs to change.
Evidence — screenshot 4
M-3

Which mailbox is actually indexed?

Needs confirmation
QuestionThe tester connected a secondary email but reports the missing message is in his primary inbox. If only the secondary account is connected, NO RESULTS may be technically correct behaviour reported as a bug — in which case M-1 becomes a messaging and scope-transparency problem rather than a retrieval failure.
ActionConfirm which accounts are connected and which are indexed before debugging retrieval. The answer changes the fix.

Scenario Analysis

Input used — Title: “What should AxioMetis work on given that 40% of the experts believe that a World war is possible in 2030” · Situation: “www.axiometis.com. I want to brainstorm our long term strategy with you.”

S-1

Probability bands are static boilerplate

Blocker
Observed
BASE 60% Constructive baseline — current trajectory holds. UPSIDE 30% Upside if catalysts align. TAIL 10% Tail: black-swan disruption.
Values and captions are generic and show no relationship to the input.
ExpectedBands derived from the described situation, with reasoning.
ImpactOutput is not usable as analysis.
Evidence — screenshots 5, 6, 7
S-4

Risk classifier contradicts the stated premise

High
ObservedAdvisory note reads: “Low risk tier — no risk markers detected in the described context. Outcome distribution favors the constructive baseline.” The described context explicitly states “40% of the experts believe that a World war is possible in 2030.”
ExpectedA scenario framed around war risk should not classify as “no risk markers detected”.
ImpactDemonstrably wrong on its own input, and the contradiction is visible to the user on screen.
Evidence — screenshots 6, 7, 8
S-2

“Key signals” are unrelated news headlines

High
Observed
• A no-brainer for protecting your brain • Who is capable of evil? • How to make AI safe—and lessen dependence on America and China • Donald Trump's blind alley • It's too darn hot. Blame global dimming
None relate to AxioMetis, strategy, or the stated premise. These are the same crawled world events that appear in Cue Cards.
ExpectedSignals selected for relevance to the scenario, or an honest “no relevant signals found”.
Evidence — screenshots 6, 7, 8
S-3

Supplied URL ignored

Medium
Observedwww.axiometis.com given as situation context; nothing indicates it was fetched or used.
ExpectedEither use it, or state that URLs are not read.
Evidence — screenshot 5

Likely common cause for S-1, S-2 and S-4

The services layer performs scenario evaluation with deterministic keyword/similarity matching over ingested world events, and the verification scorer is a stub returning fixed values. With no keyword overlap between the scenario text and the ingested corpus, it would fall back to default bands and emit whatever events are in the namespace — which matches this output precisely. Flagged as likely, not confirmed — verify before quoting.

Cue Cards

Q-2

Boilerplate scenario promoted into the brief

High
ObservedA PATTERN card carrying the scenario title and the generic advisory text verbatim as its body.
ExpectedLow-confidence or non-substantive output should not surface as an intelligence signal.
ImpactCompounds S-1 and S-4 — filler now appears in the brief surface as a finding.
Evidence — screenshot 8
Q-1

Raw internal identifiers exposed

Medium
Observed
Source: world_event:61c1ac91-66c3-462d-bc04-3fef345928df Source: pattern:a6decb5e-9a9c-4e43-a040-a9f059f7878d
ExpectedA readable source (publication, feed name) or nothing.
Evidence — screenshot 8
Q-3

No priority discrimination

Low
ObservedEvery card badged MEDIUM.
ExpectedMeaningful spread, or drop the badge.
Evidence — screenshot 8

Positive contrast worth noting

Cue Cards render titles, summaries and type badges correctly and are differentiated — which makes Correspondent (C-1) the outlier rather than a platform-wide pattern.

Principal Context

Session: Cartographer chat, Active · 15:23. Tabs read Signals 198 · Scenarios 0 · Sentinel 0. The principal's connected Gmail is on the domain axiometis.com and their company is Axiometis.

P-1

EVA does not know its own principal's employer

Blocker
ObservedPrincipal asks: “how does it impact axiometis”. EVA replies: “I don't have any information about Axiometis in memory yet. This appears to be a new entity that hasn't been tracked.” and then asks the principal: “1. What is Axiometis? (company, competitor, partner, client?) 2. What's your relationship to them? 3. What space do they operate in?”
ExpectedThe assistant knows who it works for. Asking a principal to explain their relationship to their own employer is the single most damaging output seen in this round — it directly contradicts the “sworn to one principal” premise.
ImpactEvery downstream feature that should be relevance-scoped to the principal's business has no anchor: signals, scenarios, drift, alerts.
Evidence — Cartographer chat screenshot
P-2

Email domain never used to derive organisation

High
ObservedThe connected account is @axiometis.com. That alone identifies the employer, and it is not used.
ExpectedAt minimum, seed the principal's organisation from the connected mailbox domain at connect time; ideally confirm it with the principal during onboarding.
Evidence — Cartographer chat screenshot; tester confirms the account domain
P-3

198 ingested signals, none connected to the principal's business

High
ObservedSignals 198, yet the assistant cannot relate any of them to Axiometis. Consistent with S-2, where “key signals” were unrelated general-interest headlines.
ExpectedIngestion scoped or scored against the principal's actual business context.
NoteSame underlying gap as P-1 — without knowing the principal's organisation, no relevance ranking is possible, so the crawler surfaces generic news.
Evidence — Cartographer chat screenshot; screenshots 6–8
P-4

Scenario count reads zero after a scenario was run

Medium
ObservedScenarios 0, despite the tester having run a scenario analysis that produced output and appeared in Cue Cards as a PATTERN card.
ExpectedThe counter reflects stored scenarios — or the scenario is genuinely not being persisted. Either a display bug or a persistence bug.
Evidence — Cartographer chat screenshot vs screenshots 5–8

Suggested Triage Order

  1. P-1 / P-2 — the principal's identity and organisation are the anchor everything else scores against. S-2, P-3 and much of the scenario weakness are downstream of this, so fixing relevance without it is not possible.
  2. C-3 — Correspondent's core loop is broken and the state is actively misleading.
  3. M-3 then M-1 — confirm indexing scope before debugging retrieval; the answer changes the fix.
  4. S-1 / S-4 — scenario analysis is shipping wrong output, and S-4 is visibly self-contradicting.
  5. Q-2 — stop promoting boilerplate into the brief, independent of the S-1 fix.
  6. C-1, C-4, C-5 — context and identifiers across Correspondent.
  7. P-4 — determine whether it's a counter bug or a persistence bug.
  8. C-2, Q-1, Q-3 — presentation cleanup.

Not Defects

E-1

Google “untrusted app” warning

Expected
Expected for an unverified OAuth application. Tester flagged it as anticipated. Will require Google verification before non-test users.
OK-1

Secondary email connected successfully

Working
Connection flow worked.