Captures the locked architectural decisions, module layout, and 6-commit sequence. Implementation proceeds autonomously per this plan. Refs: decisions from session 2026-09-15.
13 KiB
Agentik: Toolsets + Storage Refactor — Implementation Plan
Status: LOCKED. Implementation proceeds autonomously, no mid-implementation pings.
Date: 2026-09-15.
Resolved questions (defaults applied)
- Q1 (tool-result language): English — for consistency with the toolset section in the system prompt (also English).
- Q2 (which info in
Toolset 'X' activated.text): Q2-α minimal —"Toolset 'X' activated.", no tool names. Tool descriptions are already in the next request'stools[]array, no duplication needed.
No further open questions.
Architectural decisions (locked)
| Parameter | Decision |
|---|---|
| Layers | :agent-core (existing :standalone core), :agent-toolsets (new wrapper), :storage-core / :storage-inmemory / :storage-sqlite / :storage-android (new). |
| Default toolsets | listOf(). When empty, zero agent changes: no enable_toolset / disable_toolset tools, no toolset section in prompt. Full invisibility. |
| Activation scope | per-conversation (conv.id). |
| Dispatch outcome | Run(value: String) | Error(message: String). Substituted variant dropped. |
| Auto-activation | Stays in ToolsetDispatchPolicy only — never mentioned in the system prompt. Falls through silently when the model "goofed". |
enable_toolset / disable_toolset audit |
H2: written as ordinary ToolCall / ToolResult rows in SQLite (model sees them in its own history on subsequent turns). |
ToolsetContribution fields |
name: String + description: String + tools: List<NamedTool>. No enabledByDefault. |
| Tool naming | ${toolsetName}_${verb}; prefix = name. |
| Catalog rendering | Both ACTIVE and INACTIVE rows render as name — description, identical format. |
Toolset section language |
English. |
| Tool-result language | English (per Q1). |
| Activation timeout | 10 minutes since last use; lazy cleanup on every activeNames(convId) call. No background timers. |
| Persistence | In-memory only. Not persisted. New conversation = fresh registry. |
| Storage | StorageBundle = MessageStore + WorkingMemoryStore + ReflectionStore + SkillStore. SQLite is one impl, others possible. |
| Skills vs toolsets | Separate concepts. No link between skill index and toolset catalog. |
| Module split | A5-γ: extract interfaces, defer actual Android impl. |
Module layout (final)
agentik/
├── storage-core/ NEW (KMP)
│ ├── MessageStore / WorkingMemoryStore / ReflectionStore / SkillStore interfaces
│ └── StorageBundle aggregator
│
├── storage-inmemory/ NEW (KMP, tests)
│ ├── InMemoryMessageStore
│ ├── InMemoryWorkingMemoryStore
│ ├── InMemoryReflectionStore
│ ├── InMemorySkillStore
│ └── InMemoryStorageBundle
│
├── storage-sqlite/ NEW (JVM, refactor of existing)
│ ├── Sqldelight-backed impls of all four stores
│ └── SqliteStorageBundle
│
├── storage-android/ NEW (Android, deferred — placeholder
│ └── (placeholder file; full impl comes with Android module)
│
├── agent-toolsets/ NEW (KMP)
│ ├── ToolsetContribution (name + description + tools)
│ ├── ToolsetRegistry
│ ├── ToolsetDispatchPolicy (wraps inner DispatchPolicy)
│ ├── SystemPromptToolsetSection (SystemPromptContributor)
│ ├── EnableToolsetTool / DisableToolsetTool (NamedTool)
│ └── ToolsetWrapper — not exposed as a separate class; the registry + policy +
│ section are constructed and injected individually.
│
├── standalone/ MODIFIED
│ ├── Main.kt — wires new modules (StorageBundle + ToolsetRegistry)
│ ├── build.gradle.kts — new dependencies
│ ├── ChatAgent / ChatConversation — depends on StorageBundle (interface), not
│ │ SqliteStores directly. accept ToolsetRegistry + contributions via ctor.
│ └── all existing tests still green.
Out of scope (kept in :standalone for now): Main, transport adapters (/agentik,
/agui, /a2a), LiteLlm backend wiring, debug endpoints, MCP integration.
Commit sequence (six commits, each builds + tests green)
Commit 1 — :storage-core interfaces + bundle
New module storage-core (KMP, commonMain only):
MessageStore.kt— interface (append, getMessages, tokenStats).WorkingMemoryStore.kt— interface (append, list, compact, archive).ReflectionStore.kt— interface (insert, listRecent, count, listForConversation, deleteOlderThan).SkillStore.kt— interface (catalog, upsert, remove, exists, all).StorageBundle.kt—data class StorageBundle(val messageStore, workingMemoryStore, reflectionStore, skillStore).- One umbrella test:
StorageInterfaceContractTestasserting parameter naming is right (compile-time only).
Gradle setup:
settings.gradle.kts—include(":storage-core").storage-core/build.gradle.kts— KMPcommonMainonly withapi kotlinx-coroutines-core,api kotlinx-datetime. No JVM target yet.
Verification:
./gradlew :storage-core:build— green../gradlew :storage-core:jvmTest— green (placeholder test).
No :standalone modifications yet.
Commit 2 — :storage-inmemory impl
New module storage-inmemory (KMP, commonMain):
- All four
InMemory*implementations backed byConcurrentHashMap+MutableStateFlow-ish snapshots forgetMessages(..): Flow<MessageRecord>. InMemoryStorageBundlefactory.
Tests (KMP commonTest):
InMemoryMessageStoreTest— append + getMessages (paged flow).InMemoryWorkingMemoryStoreTest— append + list + compact + archive.InMemoryReflectionStoreTest— insert + listRecent + count.InMemorySkillStoreTest— catalog + upsert + remove.
Gradle setup:
- Depends on
:storage-core.
Verification:
./gradlew :storage-inmemory:allTests— green.
Commit 3 — :storage-sqlite refactor
New module storage-sqlite (JVM):
- Move existing
SqliteStores(and related Sqldelight code) here. - Split into
SqliteMessageStore,SqliteWorkingMemoryStore,SqliteReflectionStore,SqliteSkillStore. SqliteStorageBundle(val db: AgentikDatabase). Existing schema/migrations are unchanged. Reads from existing.sqfiles.
Migration:
- Existing tests that depend on
SqliteStorescontinue to work — keep a thin compat:SqliteStoresbecomes a deprecated alias:@Deprecated("Use SqliteStorageBundle") class SqliteStores(db: AgentikDatabase): StorageBundle by SqliteStorageBundle(db)
Tests (JVM):
- Existing sqldelight-backed tests still green.
- Add contract tests for new individual stores.
Verification:
./gradlew :storage-sqlite:jvmTest— green.- All existing
:standalonetests that referencedSqliteStoresstill compile (via the @Deprecated alias).
Commit 4 — :agent-toolsets core
New module agent-toolsets (KMP, commonMain):
ToolsetContribution.kt— data class.ToolsetRegistry.kt—class ToolsetRegistry(clock: Clock = Clock.System).activeNames(convId): Set<String>— lazy cleanup.enable(convId, name): String— 4-case response table (see A2 above).disable(convId, name): String— 4-case response table (see A3 above).DEFAULT_TIMEOUT_MS = 10 * 60 * 1000.
ToolsetDispatchPolicy.kt—interface DispatchPolicy { dispatch(call, sessionId): DispatchOutcome },class ToolsetDispatchPolicy(inner: DispatchPolicy, registry, contributions, coreTools):- Dispatch loop:
while (true): if name in resolved (core + active sets) tools: return inner.dispatch(...) if name has prefix matching known toolset T: registry.enable(sessionId, T); continue return Error("tool 'X' not found")
- Dispatch loop:
SystemPromptToolsetSection.kt—class SystemPromptToolsetSection(contributions, enabledSets)implementing:agent-core:SystemPromptContributor(defined here for now; could later live in:agent-core).EnableToolsetTool.kt/DisableToolsetTool.kt—NamedToolimplementations. Args schema:{ "name": "<string>" }as JSON.ToolsetSystemMessages.kt— companion withcoreToolsDescription: List<String>(just["enable_toolset", "disable_toolset"]).
Tests (commonTest):
ToolsetRegistryTest:- enable + activeNames immediately reflects.
- enable idempotent (returns "already active").
- disable on inactive returns "deactivated" (per A3 case 2).
- timeout cleanup via injected
Clock. - per-conversation isolation (different convIds).
ToolsetDispatchPolicyTest— syntheticpingtoolset:- model calls
ping({})without enable → auto-enabled, executed, returns "pong". - model calls
enable_toolset({})with empty args → error message. - model calls
ping({})with no such toolset registered → error.
- model calls
Commit 5 — :agent-toolsets integration (SystemPromptContributor)
Same module, adds:
- Define
SystemPromptContributorinterface inside:agent-toolsets(or move to:agent-core, but:agent-toolsetsalready has it; keep here for now). SystemPromptToolsetSectionrenders:## Toolsets Named groups of tools. One set per conversation. Use enable_toolset({"name": X}) to add a toolset; disable_toolset({"name": X}) to remove. Both are idempotent. ACTIVE name — description INACTIVE call enable_toolset({"name": X}) to add name — description ...:standalone/Main.kt— whencontributions.isNotEmpty():- Build
ToolsetRegistry(). - Add
EnableToolsetTool(registry)andDisableToolsetTool(registry)toChatAgent.tools. - Inject
SystemPromptToolsetSectioninto the system-prompt contributors chain. - Wrap the
DispatchPolicywithToolsetDispatchPolicy(...).
- Build
Tests:
SystemPromptToolsetSectionTest— renders both ACTIVE and INACTIVE rows identically.- Default (:standalone config) has
contributions = listOf()→ no tools, no prompt section.
Commit 6 — :standalone swap StorageBundle
Changes to :standalone:
Main.kt— buildSqliteStorageBundle(db)instead ofSqliteStores(db).ChatAgentconstructor:(..., stores: StorageBundle, ...)(was:SqliteStores).- All references to
SqliteStores.messageStore→stores.messageStore(etc.). build.gradle.kts— addimplementation(project(":storage-sqlite")), remove direct reliance on sqlite plumbing internals if any.
No behaviour change: existing tests still pass.
Verification:
./gradlew :standalone:jvmTest— all green (was 264 tests pre-refactor).- Build
:standalone:fatjar— works. - Smoke test against
/tmp/agentik-sandbox— e2e green (model still answers, memory still works, no regressions).
Verification at every commit
After each commit:
./gradlew :<module>:build— green.- Affected module's tests — green.
./gradlew :standalone:jvmTest— green (no regressions in the biggest test suite). Once:standalonestarts depending on the new modules in commit 5/6, this becomes the canonical regression check.- After commit 6 — run the e2e smoke probe against the sandbox (curl
/agentik/agui/a2ahealth + one turn).
Risks and mitigations
-
Risk: existing
SqliteStoresreferences scatter across:standalone.
Mitigation: keep@Deprecatedalias until all references are swept (commit 6); final sweep at commit 7 (deferred). -
Risk:
ChatAgentctor signature changes break many call sites.
Mitigation: introduceStorageBundleas a thin ctor param; existing ctors that default toSqliteStorageBundle(db)still work. -
Risk:
DispatchPolicyis currently implicit (direct call totoolsByName). Wrapping it from outside may break test doubles.
Mitigation: introduceDispatchPolicyinterface in commit 4 alongside the wrapper. Existing fakes gain the one-method interface trivially. -
Risk: toolset section length in prompt — for many toolsets, ~5 lines × N contributions.
Mitigation:descriptionfield is short (<100 tokens); limit contributions count via agent config.
Open items for future (NOT in this implementation)
- Android
:storage-androidimpl (A5-γ defers this). - Wrapper-class abstraction over (registry + policy + section) once Android needs it.
- Test-time Clock injection beyond
ToolsetRegistryTest. - Pre-validation of
descriptiontext via LLM (probably not worth it). - Exporting
ToolsetDispatchPolicyto consumers outside:standalone.
How to resume after context loss
If this session is compacted and the plan lost:
- Read this file:
agentik/docs/TOOLSETS-PLAN.md. - Verify current commit:
git log --oneline -6— should show commits in the order above. - Resume from the next commit in the sequence not yet landed.
If commits 1-3 are landed but no further, jump to commit 4. If commits 1-5 are landed, jump to commit 6.