From a method to a working product.

Phroneme connects a practical planning and reflection workflow with a small evidence library. The Fitness Lab lets a person choose activities and days, record completion, reflect on the week, and export or delete their journal.

The open-source toolkit grew from MindForge’s original six Claude methods: deep research, first principles, steelman/red-team, decision science, Socratic inquiry, and cross-domain transfer. Version 2 preserves those methods and adds a seventh fitness-evidence skill, an executable evidence contract, and the portable Fitness Lab.

The full Phroneme product and reusable public toolkit are distinct artifacts. This case study links the working experience and the public implementation of its evidence and journal components.

A narrow contract you can inspect.

The reviewed library contains three primary agency sources and five scoped paraphrases of general adult activity guidance. Source records carry attribution, a public URL, editorial review dates, status, and a declared relationship to each claim.

The validator rejects unknown claim IDs, unrelated citations, changed quantities or populations, missing review metadata, withdrawn sources, and overdue reviews. A review deadline is an editorial maintenance rule; passing that date does not mean the original guidance became false.

The contract

A reviewed wording match is a catalog result. It is not independent scientific verification or a personal recommendation.

The explorer uses complete curated statements and a small explicit phrase list, with normalization for presentation differences. It does not grade arbitrary paraphrases or call a model. Unknown wording returns insufficient-evidence: the small library cannot assess it, rather than declaring it false.

Finish the loop in the browser.

The planner validates dates, scheduled activities, completion records, and saved data. Progress is the count of completed versus planned check-ins. The printable journal escapes user-entered text, and JSON export preserves the plan and reflection.

The browser workflow stores the journal locally, offers explicit export, and removes the saved browser copy when the user deletes the plan. The lab does not send journal content to a backend or model. A shared browser profile can still expose that local copy, and exported files remain under the user’s control.

The useful engineering boundary is concrete: the software can record a plan and what the user checked off. Those records do not establish improved adherence, fitness, or health.

Test the contract. Evaluate the method separately.

The public version has 54 passing software tests: 48 evidence checks and development fixtures, plus six planner tests. They cover altered numbers and populations, unsupported source relationships, invalid review dates, hostile text, malformed storage, calendar boundaries, completion integrity, and export escaping.

Those tests establish behavior for the declared software cases. A controlled model comparison has not been run. The proposed evaluation protocol compares a fixed model with no method, a neutral checklist, and the fitness-evidence skill under identical source packets, with blinded human review.

The next research question is whether the method improves source support and uncertainty communication on held-out tasks. That question is separate from product usability and from any health outcome.

Implementation and scope are documented in the public version 2 README. The reusable toolkit is MIT licensed. This is a self-published account of a working product and its open-source contract.

Have a question or a useful counterexample?

Start a conversation