How it works

How the AI stays accurate

Stackeddd deals in specifics — your real protocols, your labs, and an AI coach that reasons over your own data. That only earns trust if the accuracy is real — so here is exactly how I keep it honest, mechanism by mechanism, including where it still falls short.

Grounded in your dataCurated knowledge baseTested before every AI update

Grounded in your data

What “grounded” actually means

Before your question ever reaches the model, the app reads your logged stack and writes it out as a block of facts. The coach answers from your numbers — not an average case stitched together from message boards.

What the coach seesassembled fresh on every message
Your active protocolsdoses, schedules, start dates, and titration steps
Your inventorypeptide and HRT vials with computed totals and runway
Your latest labseach marker flagged in- or out-of-range against your baseline
Your injection logyour logged shots and adherence
Your medssplit into currently taking vs. on hand but not taking
Your trendsrecent nutrition, sleep, and fitness
Cycle & side effectscycle status and anything you have recently flagged
Pinned factsthe notes you have told it to always remember

This is the single most important thing for accuracy: the model can only be accurate about data it can see. Most of the real accuracy gains in Stackeddd’s history came from improving this data block — not from clever wording. On top of it, the app pulls in reference material on demand: your question is matched against the knowledge base by meaning and by keyword, and only a handful of the most relevant entries ride along. So the coach is reasoning from your real protocol plus vetted reference, every time.

1
Your data becomes text
A function on your phone reads your logged stack and writes it out as a structured block of facts — rebuilt fresh on every single message.
2
Reference is retrieved
Your question is matched against the knowledge base by meaning and by keyword; only the few most relevant entries are attached.
3
The request is assembled
The standing rules, your data block, those reference entries, and your recent chat are combined into one request.
4
The answer streams back
The model replies word-by-word, reasoning from your numbers — not a forum-average case.
The math isn’t AI at all

The blood-level chart, the dose and blend calculators, and the adherence math are ordinary code, not the model. Same input, same output, every time — there is nothing there to hallucinate. Anything that can be deterministic is.

Anti-fabrication

What happens when it doesn’t know

The failure that matters most for a tool like this is confident invention — the coach stating something about your data that your data doesn’t support. The standing rules draw a hard line around it:

  • It works only from the facts in your data block. If something isn’t there, it is told to say so rather than guess a number.
  • When a figure can be worked out from your data, it recomputes it rather than repeating a stale answer from earlier in the conversation.
  • The line is drawn specifically around your data. General, well-established information is fair game; making up your lab value, dose, or inventory total is not.

None of this is a promise of perfection. It is a set of rules that make fabrication rare — and, just as importantly, a test suite that measures how rare, which is the next section.

The knowledge base

How the reference is built

Every entry is synthesized into original prose — written up from primary literature and the real discussion in this community, then reviewed and edited by me before it ships. Nothing is copied or redistributed word-for-word, and I decide what makes it in. That is a deliberate line: I would rather ship less reference I can stand behind than more I can’t.

  • Harm-reduction framing. Peptide and compound profiles written to inform, never to encourage a dose or a protocol.
  • Original synthesis, not copied. Written up from primary literature and community discussion, then reviewed — nothing is redistributed word-for-word. A legal and quality decision, not a shortcut.
  • Curated, not crowd-sourced. You can’t add your own entries — you suggest a topic through the app, and I decide what gets added.

Retrieval quality — whether the right entry actually surfaces for a given question — is guarded by its own deterministic test suite that runs alongside the rest.

The test suite

How the AI is tested

This is the part almost no one shows. Every change to the coach’s instructions runs against a suite of frozen, adversarial scenarios before it can ship — so a new rule can’t quietly break an old one.

~60
frozen scenarios re-checked on every run
attempts per scenario — it must pass at least two
Every
update to the prompt is gated on a full run first

Each test is a real trap

A scenario is a complete synthetic phone — profile, labs, protocols, inventory — chosen to bait a specific mistake (say, a borderline-low HDL loaded before a question about a compound that can move it). It is built into the same request the app sends, and the answer is graded. Many are pinned to a real incident — a bad answer that happened once, frozen so it can never quietly come back — and the rest are proactive traps, written to catch a mistake before it happens.

What it’s tested against

The traps map to the specific ways this kind of AI goes wrong:

FabricationMade-up numbersInventing a value, or pairing the right number with the wrong date or dose.Caught by: Canonical-data rules, audited by a numeric trace in eval runs.
False reassuranceUnearned comfort“That won’t affect your lipids” — volunteering safety it can’t actually promise.Caught by: No-false-reassurance rules, watched by a dedicated rate canary.
SofteningDownplayed labsCalling an out-of-range marker “nothing to worry about.”Caught by: Lab-interpretation rules that keep flags flagged.
SycophancyJust agreeingEndorsing whatever you propose because you proposed it.Caught by: Anti-sycophancy rules that require an honest read.
Missed flagIgnored contextAnswering a symptom question while ignoring a relevant med or condition you logged.Caught by: A symptom sweep that forces the relevant flags into view.
Off-targetAnswering the wrong thingA three-bullet essay for a one-number question — or missing what you actually asked.Caught by: Response-depth and answer-what-was-asked rules.

How each answer is graded

Up to three layers grade a response, in order of how much they are trusted:

1
Exact text checks— the gate
Deterministic patterns for banned phrases and required content — the recomputed number, the disclaimer. Dumb, literal, and the only layer that can actually fail a test and block a release.
2
A second model grades the answer— advisory
A separate AI scores each response against a tight rubric with the ground truth supplied. Only a catastrophic score flags it — a model grading a model is noisy, so it warns rather than gates.
3
Every cited number is traced— opt-in audit
An audit I can run over an eval batch — it pulls out every figure and date an answer cites and checks each against the data the model was given: grounded, correctly derived, contradicted, or unsupported. Off by default; used to spot-check, not to gate.

The honest part: some failures can’t be zeroed

AI responses are stochastic — the same question can get different answers. Rules make a failure rare; they cannot make it impossible. Even the most-reinforced rule still slips roughly 5–10% of the time. Pretending otherwise would be the overclaim.

So the touchiest rule — the one against false reassurance — isn’t graded pass/fail. Its scenario runs 20 times and reports how often the rule slipped. If that rate climbs past a set threshold, well clear of the normal floor, it alarms — an early warning that a recent change eroded the rule. That is the difference between hoping the AI behaves and measuring whether it does.

What the tests can’t see

Being straight about the limits: the suite doesn’t grade photo reading, and it can’t catch a brand-new kind of mistake nobody has thought to write a test for yet. Those get found the hard way — in real use — and turned into the next frozen scenario. New failures come from real chats; the suite’s job is making sure fixed ones stay fixed.

The honest limits

What this is — and isn’t

Stackeddd is a tracking and research tool, not a doctor. The coach doesn’t diagnose, treat, or prescribe, and everything here is educational. It is built to help you understand your own data and ask better questions — not to replace the person who manages your care.

  • Verify before you change anything. Any decision to start, stop, or adjust a dose or protocol belongs with your prescribing physician. Take the coach’s read as a prompt to ask, not as an instruction to act.
  • It can still be wrong. The rules and tests make mistakes rare, not impossible. If a number looks off, trust your own records — and tell me, so it becomes the next test.
  • Your data stays yours. No account, and no analytics beyond anonymous crash reports. Your logged data lives on your phone; when you use an AI feature, only the context needed passes through a relay that holds the API key and keeps nothing. Stackeddd runs no server that stores your data.

Stackeddd

Track it instead of guessing.

Your stack, your labs, and an AI coach that reasons from your own data — free to start, no account required.