- proto: Agent.info (AgentInfo: name/description/usefulness) — человекочитаемое
имя отдельно от opaque id; сериализация и тесты
- server: GET {path} отдаёт Agent.info; GET /conversations/{id}/record —
ConversationRecord без handle'а (для клиентского кэша после AgentEvent.Created)
- outbox: Event.Working — первый event хода, эмитится из Conversation.send()
до LLM-цикла и turnLock, чтобы UI показал спиннер сразу
- standalone: AGENTIK_NAME/AGENTIK_DESCRIPTION/AGENTIK_USEFULNESS → AgentInfo
- client: AgentClient.create eagerly фетчит info (GET {baseUrl})
- client: HttpConversationStore.get читает /record (ConversationRecord,
а не ConversationSnapshot — рассинхрон типов)
- client: noReadTimeout() на POST /conversations/{id}/messages — сервер отвечает
по завершении всего хода агента (реально 0.5–144 с), дефолтные 15 с рвали
живую реплику на клиенте
- journal-api: ConversationRecord @Serializable
- ksqlite 0.1.3 → 0.1.4
- .gitignore: runtime-данные standalone-агента и hs_err-дампы
ksqlite 0.1.3 опубликован в Maven Central — POM/module-metadata/jar все
на месте. Поднят в шести модулях (vector-index-ksqlite, journal-ksqlite,
context-ksqlite, reflection-ksqlite, standalone, memory-md-vector);
комментарии про 0.1.2 в build.gradle.kts обновлены.
Расширение :journal-api: добавлены два count-метода (total + count after
cursor) в интерфейс JournalStore — реализованы в :journal-ksqlite /
:journal-inmemory. Над ними добавлены HTTP-endpoint'ы
GET /journal/conversations/{id}/count
GET /journal/conversations/{id}/count?after=
(объединены в один маршрут с опциональным параметром) и HTTP-клиент
HttpJournalStore.count/concount. Сервер-фасад расширен тестом
JournalRoutesCountTest (5 кейсов через embedded CIO + реальный
InMemoryJournalStore).
:server:jvmTest 10/0, :standalone:jvmTest 129/0, :journal-ksqlite:jvmTest 25/0,
:journal-inmemory:jvmTest 19/0. jvmTest агрегат 425/0/0.
Adds `agent.conversationStore` (read-only view on `conversation` table) to
the :proto Agent interface, plus `agent.renameConversation(id, title?)`
command. Client-side cache in :client is built from a snapshot
(`remote.listFlow(0)` → `local.upsert(...)`) + live updates via
`outbox.agentEvents()` (Created/Deleted/Renamed/Touched).
Changes:
- :journal-api — split `ConversationStore` (read-only: get/list) and
`MutableConversationStore` (CRUD: upsert/delete/rename/touch);
`ConversationStore` gained `listFlow` (cold-flow paging via `list`).
- :outbox-api — `AgentEvent.Touched(date, id, updatedAt)` event so
client cache stays fresh after `send()` (which bumps `updatedAt`).
- :proto.Agent — added `conversationStore: ConversationStore` property,
added `renameConversation(id, title?): Instant?` command, removed
`getConversations(offset, limit)` (now: `conversationStore.list(...)`).
- :server — `GET /conversations` now returns `List<ConversationRecord>`
(lightweight metadata, no handle/image-support flags); `PATCH
/conversations/{id}` uses `agent.renameConversation` and returns
the updated `ConversationRecord`.
- :journal-inmemory — expanded targets to jvm+macos+linux+mingw (matches
:client); moved `InMemoryMutableConversationStore` here from
:storage-inmemory so :client can use it without pulling ios targets.
- :storage-inmemory — depends on :journal-inmemory.
- :storage-ksqlite — pre-staged rename `KsqliteConversationStore` →
`KsqliteMutableConversationStore` to match the new interface split.
- :standalone — `ChatAgent` exposes `conversationStore` as a read-only
view of its `mutableConversationStore`; emits `AgentEvent.Touched`
after each `send()` (after `conversationStore.touch(id, ts)`).
- :client — new `HttpConversationStore` (read-only HTTP impl);
`AgentikAgent` wraps the agent with `wrapWithLocalConversationCache`
so the client sees an in-memory cache (snapshot + outbox events)
instead of direct HTTP. Cache scope + HttpClient + background job
all cancelled in `agent.close()`.
- :client/README — new «Кэш списка бесед» section with the
`listFlow → upsert` / `agentEvents → apply` pattern and a note that
`conversationStore` is read-only (writes only via Agent commands).
All 96 jvmTest tasks green.
Three protocol-level changes from Android-client review (items 1-3, 5-6):
1) toolName denormalization in ToolResult (3 layers):
- :outbox-api/Event.ToolResult: +toolName: String? = null
- :journal-api/MessageRecord.ToolResult: +toolName: String? = null
- :proto/Message.ToolResult: +toolName: String? = null
- :storage-ksqlite, :journal-ksqlite ResultPayload codec: +toolName
- :standalone/ToolDispatcher, ConversationLoop: thread toolName = call.name
Nullable + default = backward-compat for already-persisted histories
and existing clients.
2) Drop proto/Event.kt, AgentEvent.kt, CommonEvent.kt typealiases.
is proto.Event.End failed with 'Unresolved reference End' (alias
loses nested-class access). Use pw.binom.agentik.outbox.{Event,
AgentEvent, CommonEvent} directly everywhere — :proto already has
api(:outbox-api), the package is visible to consumers, no shim
needed. 21 files rewired, 3 files deleted.
3) Rename Event.ToolResult.id → toolCallId (option B per user).
In :outbox-api Event.ToolResult.id == Event.ToolCall.id (one value,
one name); the persistent journal keeps MessageRecord.ToolResult.id
as its own PK + toolCallId as FK to the call — different semantics,
left untouched. Fixed ToolDispatcher bug: emitted id = resultId
while KDoc claimed id == ToolCall.id; now emits toolCallId = callId.
4) Remove Conversation.events() from :proto; OutboxStore is sole event source.
Conversation is a pure per-conversation abstraction (send/getMessages/
rename/close). Live events only via agent.outbox.conversationEvents/
agentEvents/events. HTTP route /conversations/{id}/events stays for
wire-compat but routes through outbox internally (map { it.event }).
jvmTest green (95 tasks).
- :memory-md-vector (KMP jvm+linuxX64+mingwX64): .md-файлы как source of
truth, векторный индекс (sqlite-vec) как derived cache. reconcile()
на старте: orphan-cleanup + content-hash-gated re-embed. Гибридный
скор 0.7*vector + 0.3*keyword. Заменяет EmbeddingProvider на
KMP-TextEmbeddingExecutor из :memory-api.
- :reflection-api: новый 4-й API-модуль (Reflection, ReflectionStore,
ReflectionEvent). Зависит только от :memory-api.
- :journal-api получил ConversationRecord/ConversationStore/Ids (бывший
:message-store-api, полностью удалён). :outbox-api получил Event,
CommonEvent, AgentEvent (бывший :event-store).
- :memory-api получил MemoryVectorIndex + NoteMatches +
TextEmbeddingExecutor (suspend-обёртка над TextEmbeddingExtractor).
- :memory-vector KMP-цели достигнуты через commonMain-only TextEmbedding-
Executor, EmbeddingProvider выпилен; :memory-md-vector тянет
text-embedding-api транзитивно через :memory-api.
- :standalone flatten в commonMain/commonTest завершён (тесты из jvmTest
переехали в commonTest). Включён optional деп :memory-md-vector через
AGENTIK_MEMORY_BACKEND=md-vector.
jvmTest: 96 задач, 407 тестов, 0 падений.
ksqlite 0.1.2 опубликован в Maven Central (был только в локальном ~/.m2);
text-embedding-kmp v4 — в caffeine Nexus (поддержка нативных целей):
jvm, android, linuxX64/Arm64, macosX64/Arm64, iosX64/Arm64/SimulatorArm64.
Без этих апдейтов CI release-пайплайн падал с unresolved-dependencies
на любом свежем коммите после acc7237e5 (введение ksqlite).
Дополнительно: игнорируем локальный opencode config.json.
- Replaced `EmbeddingProvider` with cross-platform `TextEmbeddingExecutor` for native target compatibility.
- Introduced `:memory-md-vector` module combining vector-cache and `.md` file-based memory systems (`hybrid` backend).
- Updated `SiglipEmbeddingProvider` to use KMP `TextEmbeddingExtractor` and streamlined compatibility via `asExecutor`.
- Added hybrid memory backend to `standalone`, supporting `.md` reconciliation with vector-cache for semantic
- Removed `:message-store-api` module and associated classes (ConversationStore, ReflectionStore, Ids, etc.).
- Migrated reusable components to `:journal-api` (conversation-related) and `:reflection-api` (reflection-related).
- Updated imports and module dependencies across all projects to reflect new structure.
- Adjusted build scripts and tests for compatibility with the new APIs.
- Replaced usages of `:message-store-api` and `:working-memory-api` with `:journal-api`, `:outbox-api`, and `:context-api`.
- Deprecated legacy `EventStore` and `MessageStore` interfaces, added `typealias` for backward compatibility.
- Updated imports across all modules with references to `:journal-api` and `:outbox-api`.
- Introduced `journalRoutes` and `outboxRoutes` in `:server` for audit log and live event stream endpoints.
- Adjusted `Agent` to expose read-only `journal` and `outbox` stores for improved modularity and clarity.
- Removed legacy Event and AgentEvent definitions from `:proto`, migrated to `:outbox-api`.
- Storage-related modules have been updated to support the new APIs consistently.
- Introduced `MutableMessageStore` for producers with an `append` operation, separate from read-only `MessageStore`.
- Updated all consumers and implementations to use the appropriate interface (`read-only` for observers, `mutable` for producers).
- Improves modularity and ensures compile-time guarantees against unintended write operations in the audit log.
- Remove legacy pw.binom.agentik.messageStore.events.EventStore (EventRecord,
EventType) and all three impls (in-memory, sqlite, ksqlite) + tests + .sq
- Drop :storage-bundle module entirely; ChatAgent / ChatConversation /
ConversationLoop / DebugRoutes now take stores individually
(conversationStore, messageStore, workingMemoryStore, reflectionStore,
eventStore) instead of StorageBundle
- Delete server endpoints /events/replay and /conversations/{id}/events/replay;
Route.agentikAgent no longer takes eventStore param
- Add :client/HttpEventStore implementing :event-store/EventStore over HTTP:
events() -> GET /events/all, agentEvents() -> GET /events,
conversationEvents(convId) -> GET /conversations/{id}/events;
exposed via AgentClient.eventStore
- :event-store: add macosX64/macosArm64/linuxArm64 targets to match :client KMP
- :working-memory-api: drop api dep on :message-store-api (no longer needed)
- :storage-{inmemory,sqlite,ksqlite}: drop deps on :storage-bundle
Выделяет append-only message log в отдельный KMP-модуль.
Цель — разделить ДВЕ сущности по своей природе:
:message-log-api — append-only audit log (User/Assistant/ToolCall/
ToolResult/Error). Никаких update, только insert + read.
Это иммутабельная история диалога.
:working-memory-api — mutable runtime context (compact, summary, WM order).
Live state. Compaction-логика.
Раньше оба жили в :message-store-api, что:
- смешивало контракты: append-only audit vs mutable runtime;
- делало невозможным лёгкого клиента который читает только audit log
без WM-runtime зависимости;
- затрудняло compaction-логике жить в одном модуле с audit-записью.
Миграция:
- В :message-log-api переехали: Content, MessageRecord, MessageStore,
MessageContext (с MessageOrigin), MessageEvent, TokenStats, TurnTokens,
helpers (encode/decodeBodyPayload, MessageBodyPayload, BodyDecoded).
Пакет pw.binom.agentik.messageLog.
- В :message-store-api остались: ConversationStore, ConversationRecord,
ReflectionStore, Ids, legacy events.EventStore (paginated replay).
Пакет pw.binom.agentik.messageStore.
- :working-memory-api: обновил deps (api → :message-log-api для Content/MessageContext).
- 23 consumer-файла обновлены (FQN renames).
- storage-sqlite/ksqlite: убраны недостижимые ветки Summary/System
(эти synthetic records живут ТОЛЬКО в :working-memory-api, не попадают
в audit log :message-log-api).
Файлы:
+ :message-log-api (5 файлов, ~280 строк)
- :message-store-api (5 файлов, ~430 строк)
~ 23 файла обновлены
Совместимость схем не меняется. Все 5 storage impl'ов (3 backend × 5 store)
работают на тех же таблицах.
Разделяет интерфейс на read-only (EventStore) и write (MutableEventStore).
EventStore (read-only, для consumer'ов):
- events(after: Instant?): Flow<CommonEvent>
- earliestEventDate(): Instant
- close()
MutableEventStore : EventStore (для producer'ов):
- + append(event: CommonEvent)
- suspend, не идемпотентный, может быть silently evicted
Зачем:
- Consumer'ы (server SSE, admin dashboard, parent agents) принимают
EventStore — compile-time гарантия что они не могут писать в store.
- Producer'ы (ChatAgent, sub-agents, A2A-bridge) принимают MutableEventStore.
- Тесты могут использовать EventStore без опасности случайной модификации.
Миграция:
- :event-store пока без implementations, поэтому ничего не сломалось.
- Когда добавим InMemoryEventStore — он будет реализовывать оба
(MutableEventStore = EventStore + append). Подписки получают только
read-only projection через приведение типа.
Также: импорт обновлён AllEvent → CommonEvent (по rename в :proto).
Разделяет монолитный :storage-core на 3 модуля с чёткими границами:
:message-store-api — MessageStore, ReflectionStore, EventStore, ConversationStore +
Content, Payload, MessageContext, Ids, MessageEvent
(audit log + event stream)
:working-memory-api — WorkingMemoryStore + WorkingMemoryEntry
(runtime context с compaction)
:storage-bundle — StorageBundle агрегатор, зависит от обоих
(только для server-side runtime)
Пакеты:
pw.binom.agentik.storage.* → УДАЛЕНО
pw.binom.agentik.messageStore.* — append-only API
pw.binom.agentik.messageStore.events.* — EventStore + EventRecord
pw.binom.agentik.workingMemory.* — WM API
pw.binom.agentik.storageBundle.* — aggregator
Зачем:
- Тонкий клиент может подтянуть ТОЛЬКО :message-store-api (~15KB, нет
compaction-логики, нет MessageStore+WorkingMemoryStore cross-deps).
- Android-agent в будущем подключит :message-store-api для audit log,
серверный runtime — :storage-bundle со всем.
- Компиляционные границы защищают от случайной зависимости от WM
в read-only клиентах (раньше один :storage-core не давал такой
гарантии).
Миграция:
- Имплементации (:storage-inmemory, :storage-sqlite, :storage-ksqlite)
обновили package + добавили deps на оба API модуля + :storage-bundle.
- Тесты из :storage-core (PersistenceTest, SqliteStoresMigrationTest,
TokenStatsTest) переехали в :standalone, получили testImplementation
на оба API модуля и импорты новых типов.
- 52 файла в :standalone, :agent-toolsets, :llm-tools, :server, :client,
:agentik-cli обновили FQN.
- :storage-core удалён.
Совместимость схем не меняется — все 5 impl'ов (3 backend × 5 store) хранят
данные в тех же таблицах, миграция между Sqlite и Ksqlite возможна через SQL dump.
Тесты:
standalone 178 ✅
agent-toolsets 36 ✅
storage-inmemory 47 ✅
storage-sqlite 17 ✅ (включая переехавшие persistence/* + tokenStats)
storage-ksqlite 36 ✅
---
Total: 314 tests, 0 failures
Добавлена собственная авторизация по токену. Это ОТДЕЛЬНАЯ подсистема:
библиотека A2A (pw.binom.a2a) имеет свой независимый token, общих типов
и общей логики не вводится.
Поведение по умолчанию не меняется: token = null -> авторизация выключена,
сервер открыт (обратная совместимость), CLI/TUI не затронуты.
Сервер (:server):
- новый route-scoped плагин BearerTokenPlugin (BearerTokenConfig);
- agentikAgent(agent, path, token) ставит плагин на всё поддерево /agentik,
когда token != null; иначе плагин не устанавливается;
- при несовпадении заголовка Authorization: Bearer <token> -> 401 Unauthorized;
- /health всегда открыт (liveness для балансировщика).
Клиент (:client):
- defaultAgentikHttpClient(token) навешивает Authorization: Bearer <token>
через DefaultRequest на весь HttpClient -> накрывает все 10 вызовов и оба SSE;
- AgentikAgent(id, baseUrl, token, httpClient) — token необязательный,
9 существующих мест создания агента не тронуты.
Standalone:
- AgentSection.authToken (env AGENTIK_TOKEN) -> /agentik;
- AgentSection.a2aToken (env AGENTIK_A2A_TOKEN) -> /a2a;
- два независимых значения, связи между ними нет.
Тесты: BearerTokenTest (5), BearerHeaderTest (3) — 401 без токена и с чужим,
200 с верным, /health открыт, null -> открыто. Мутационная проверка пройдена.
- ci.yml/release.yml: LANG/LC_ALL=C.UTF-8 — иначе Kotlin-компилятор падает
с InvalidPathException на именах тестов с типографским тире (LANG=C → ASCII)
- имена тестов: типографское тире U+2014 заменено на ASCII-дефис (16 шт)
- release.yml переведён на общий composite-action subochev/devops/publish@main
(как у asr-kmp/litert-kmp/embedder-kmp); версия = имя тега релиза
Standalone refactor — modularity + correctness improvements after
STANDALONE-REVIEW findings. Touches ~30 files. Build green, 178 tests pass.
(1) Module extractions — generic components out of :standalone:
• :llm-tools (new KMP module, package pw.binom.agentik.llm.tools)
- LlmReflector, SkillMiner, LlmMemoryReviewer, LiteLlmContextCompactor
- Parsers: ReflectionParser, SkillMiningParser, ReviewDecisionParser
- Prompts: ReflectionPrompts, SkillMiningPrompts, ReviewPrompts
• :mcp-bridge (new JVM module, package pw.binom.agentik.mcp.bridge)
- McpConfig, McpRegistry, McpLiteToolAdapter
• NamedTool moved from :standalone to :agent-toolsets/commonMain
- Generic (name + LiteTool) wrapper, used by both :mcp-bridge
and :standalone's tool dispatcher
:standalone loses ~1400 lines, depends on the two new modules.
(2) Background work → event-driven (no more interval-polling):
• New :standalone/agent/BackgroundEvents.kt — internal event bus:
- ToolCallEvent.Succeeded/Failed (emitted by ToolDispatcher after invoke)
- CompactionEvent.Triggered (emitted by CompactionCoordinator pre-delete)
- ConversationLifecycleEvent.Closing (emitted by ConversationLoop.close)
• BackgroundScheduler rewritten as event subscriber:
- On Closing: final reflection + skill mining (last-chance extraction)
- On Compaction (turnsToDelete > 10): skill mining (debounced 60s)
- On ToolFailure x2 in 60s window: reflection (debounced 5min)
- Dropped: maybeScheduleReview/Reflection/SkillMining (interval-based)
- Dropped config: memoryReviewInterval, reflectionInterval, skillMiningInterval
• ToolDispatcher emits ToolCallEvent after each invoke.
• CompactionCoordinator emits CompactionEvent before workingMemory.compact().
• ConversationLoop.close() emits Closing BEFORE agentScope.cancel() so the
subscription gets to run final reflection/mining.
Net effect: typical 30-turn conversation runs ~38 LLM calls (was: 30 main +
3 review + 3 reflection + 2 mining). With event-driven, review/mining only fire
when their triggers actually make sense (compaction about to delete, or
conversation closing).
(3) AppConfig single source of truth:
• Replaces AgentikConfig + LlmConfig.fromEnv + McpConfig.fromEnv with one
AppConfig.fromEnv() that reads all ~25 env vars in a single pass.
• Sections: AgentSection, LlmSection, McpSection, MemorySection,
EmbeddingSection, ReflectionSection, SkillMiningSection, DebugSection.
• OPENAI_CONTEXT_WINDOW / AGENTIK_GOOGLE_CONTEXT_WINDOW no longer
read twice (was a bug per STANDALONE-REVIEW E3).
(4) Other fixes inherited from earlier waves:
• Hardening — size caps on user-input boundaries:
MAX_MEMORY_CONTENT_LEN=32KB, MAX_SKILL_BODY_LEN=64KB,
MAX_MCP_CONFIG_BYTES=1MB, MAX_A2A_REPLY_LEN=10MB, MAX_PORT=65535,
blank-rejection in LlmConfig.requireEnv, URL/command validation.
• Single scope — :standalone/agent/ConversationLoop has one
agentScope (was: scope + backgroundScope).
• liteConvRef race fix — capture-then-use pattern replaces !!-after-read;
close() + runTurn.finally race on LiteConv JNI handled via
AtomicReference.getAndSet.
• SkillMiner.maxTurns / LlmReflector.maxTurns exposed as public (needed
by BackgroundScheduler for prompt sizing).
• Tests: MemoryWiringTest updated for new compaction-triggered review
behavior; all parser/test imports updated for new packages.
Test results: 178/178 in :standalone, 36/36 in :agent-toolsets — all green.
Every subproject now has README.md:
- 3 runnable modules (:standalone, :agentik-cli, :agentik-tui):
quickstart, env table, parameters, known limits
- 11 library modules: what it is, which problem solves, how to
wire it in, where versions live
Root README.md is the navigation hub (Quickstart, Modules table,
publish + CI/CD notes).
Also: ci.yml prunes the :memory-vector -x excludes now that
text-embedding-kmp artifacts are published to caffeine.
518 tests green.
Verified publish pipeline: :proto:publish to caffeine produces
pom.module + per-target klibs + sources for all 9 KMP targets.
🤖 Generated with [opencode]
- README.md в каждом подмодуле: для библиотек — описание проблемы,
подключение через maven-central/caffeine, версии в gradle/libs.versions.toml.
Для запускаемых модулей — команды запуска + переменные среды с дефолтами.
- Корневой README.md переписан как навигационный хаб: что это, где клиенты,
где серверы, как собрать, как опубликовать.
- build.gradle.kts: per-module POM-description через единую карту в rootProject.extra
(порядок важен — нужно ДО apply плагина KMP, поэтому beforeEvaluate в subprojects).
- .gitea/workflows/ci.yml (новый): build + jvmTest + shadowJar на PR/push main.
- .gitea/workflows/release.yml (обновлён): публикует библиотеки в caffeine
Nexus + собирает 3 fatjar'а и крепит их к release как бинарные ассеты.
Добавляет ModelDownloader (HTTP с Range/докачкой, опциональной SHA-256 проверкой)
и два сценария запуска скачивания встроенной модели gemma-4-E2B-it.litertlm:
java -jar agentik.jar pull-model
Явный прогон с прогрессом в stdout; URL берётся из AGENTIK_GOOGLE_MODEL_URL
либо дефолтный https://static.binom.pw/models/gemma-4-E2B-it.litertlm.
AGENTIK_AUTO_DOWNLOAD_MODEL=1 java -jar agentik.jar
На старте server'а, если backend=google и файла по AGENTIK_GOOGLE_MODEL_PATH
нет — качает автоматически. Без флага — exit 2 с понятным сообщением и
подсказкой вызвать pull-model.
Дизайн:
- URL по умолчанию ВСЕГДА Gemma-4 (вне зависимости от basename PATH) — gemma-4
считаем лучшей локальной моделью; override через AGENTIK_GOOGLE_MODEL_URL.
- SHA-256 проверка через опциональный AGENTIK_GOOGLE_MODEL_SHA256_URL.
- Resume: HEAD → если есть .part и Accept-Ranges=bytes → GET с Range: bytes=N-,
иначе restart с нуля.
- Прогресс каждые ~8 MB, финальный rename через Files.move(ATOMIC_MOVE).
Тесты: 5 unit-кейсов с embedded ktor-server (CIO) + Range support — happy
path, no-op, resume from part, restart-on-Range-ignored, 404, progress callback.
Документация: новый раздел §18 в MANUAL-TESTS.md (subcommand, auto-trigger,
resume, override URL, SHA-256 verify).
178/178 tests green.
Три фикса в runTurn/interrupt:
1. **Race condition в finally-блоке.** Раньше сбрасывал
interrupted.set(false) только если флаг был установлен при чтении
wasInterrupted в начале finally. Если interrupt() приходил между
этими двумя точками — флаг оставался true и следующий turn видел
wasInterruptedAtEntry=true → сразу short-circuit'ил без вызова LLM.
Теперь всегда сбрасываем (compareAndSet атомарен, гарантирует
следующий turn чистый).
2. **interrupt() отравлял следующий send.** Если вызывали interrupt()
в пустоту (нет активного turn'а — флаг всё равно ставился → следующий
send сразу short-circuit'ил, пользователь не получал ответа на
своё 'Ок.' после явного cancel). Теперь interrupt() проверяет
activeTurn?.isActive и при отсутствии активного turn'а — no-op.
3. **Bounded background scope для review/reflection/skill-mining.**
OpenAiLlm.send() использует runBlocking — если запустить 30+
параллельных review (по одному на беседу), IO-thread pool
голодает и ассистент висит. Вынес в отдельный scope с
Dispatchers.IO.limitedParallelism(4) — не больше 4 sync LLM
вызовов одновременно.
Тест 27/27 (см. /tmp/test-interrupt.py и /tmp/run-manual-tests.py).
После addToolResult (например memory_save result) LiteRT-LM (stateful)
возвращает дельту с финальным текстом модели. Но OpenAI-бэкенд
(stateless, litert-openai) просто дописывает tool-result в history и
возвращает пустую дельту — следующий ответ модели приходит только
при следующем send.
Без этого фикса ассистент после tool-call'а выдавал пустой текст
"\n\n" (например после memory_save).
Что меняется:
- runTurn: после addToolResult вызываем sendStreamContents с пустым
placeholder'ом (" "), который для OpenAI триггерит continuation,
а для LiteRT-LM просто даёт no-op-ответ (соберём, отбросим).
- tool_calls из continuation НЕ обрабатываем в текущем inner-while —
кладём в pendingPostToolCalls и обрабатываем на следующей outer
итерации. Иначе можно попасть в бесконечный tool-loop (fake
LiteLlm-тесты это показывают).
- emptyList() нельзя — LiteMessage требует непустой contents, поэтому
используем пробел как placeholder.
Тесты:
- tool-call loop test: toolCallCount == 2 (user send + post-tool continuation)
- live e2e на удалённой машине (192.168.76.166) с OpenAI vLLM бэкендом:
- простая арифметика (12+34=46) ✓
- memory_save + recall в той же беседе ✓
- memory persists across conversations ✓
- прерывание mid-task (генерация рассказа про космос) → partial assistant
+ ToolExchange в working memory ✓
- SSE events: start_reasoning, start_response, append_text, end ✓
Total: 341/341 green.
Radical redesign of interrupt semantics (plan: docs/TOOLSETS-PLAN.md,
phase commit 7):
1. Storage (:storage-core + :storage-sqlite + :storage-inmemory):
add WorkingMemoryEntry.ToolExchange(toolName, toolArgsJson, resultText,
wasCancelled) — one row per tool-call. Survives restarts.
2. ChatConversation:
- new fields: interrupted (AtomicBoolean), currentToolJob (Job?)
- interrupt() теперь только сигнал: ставит флаг, cancel LiteConv +
cancel currentToolJob. НЕ cancel activeTurn — пусть runTurn finally
отработает.
- runTurn обёрнут в try/finally: даже при CancellationException (от
LiteConv.cancel()) и при early-return (interrupt до старта LLM) —
finally закрывает LiteConv и эмитит Interrupted (если была отмена) + End.
- runToolAndPersist возвращает WorkingMemoryEntry.ToolExchange вместо
Pair(callId, resultText); инструмент запускается в scope.async, его
Job = currentToolJob, cooperative cancellation через Job.cancel.
Если инструмент броает CancellationException/InterruptedException →
resultText = '[cancelled by user]', wasCancelled = true.
3. GetOrCreateLiteConversation теперь мапит ToolExchange →
LiteMessage(TOOL, ToolResult, name, response) в initialMessages —
при следующем send() LLM видит честный результат вызова tool'а
через LiteRT-LM (callId не требуется, матчится по name).
4. LiteConv lifecycle: создаётся новый на каждом turn (close+recreate
семантика). Это ~2s prefill на Gemma-4-E2B, но гарантирует полную
предсказуемость: нет рекурсивных cancel-drain'ов, KV-cache всегда
консистентен с WM.
5. Тесты:
- multi-turn: 2 LiteConv-а (один на turn)
- interrupt mid-slow-stream: пустой assistant в WM, только user, события
Interrupted + End.
- interrupt after-tool: ToolExchange в WM (result=echo output, wasCancelled=false),
ToolCall + ToolResult в audit.
Total: 341/341 green.
System prompt is now built fresh at conversation create/load time
(in `buildSystemPrompt` capturing current SOUL/skills/toolsets/reflections)
and passed into LiteConversationConfig.systemInstruction. It is NOT
written to working_memory anymore.
Why: ChatAgent was freezing the system prompt into a WorkingMemoryEntry.System
row at createConversation, then reading it back on every getOrCreateLiteConversation.
This meant changing SOUL, activating toolsets, adding skills or new
reflections between agent restarts did not propagate to existing conversations
without re-running createConversation.
Fix:
- ChatAgent.createConversation: dropped the workingMemoryStore.append(System(...))
- ChatConversation.getOrCreateLiteConversation: replaces the WM-based lookup with
the in-memory systemPrompt field directly
- ChatConversation.compactPreTurn: same simplification — compaction operates only
on User/Assistant rows (plus future Summary rows); system prompt is excluded
Migration: none. Old DBs may contain dead System rows from prior versions — they
are simply ignored by the new lookup, and compaction never reads them.
Tests: 340/340 green. Updated 7 tests across ChatAgentTest + MemoryWiringTest
that asserted the old System-in-working-memory contract; they now verify the
system prompt via LiteConversationConfig.systemInstruction (what LLM actually sees).
E2E verified: 0 system rows in working_memory across all conversations,
multi-turn history reconstructs correctly after agent restart with the updated
in-memory system prompt.
Удаляю DemoToolset.kt, поле AgentikConfig.demoToolset и парсер
AGENTIK_DEMO_TOOLSET — после ночной e2e-проверки toolsets
они больше не нужны в проде. Механика покрыта ChatAgentToolsetsTest
через stub-тулы inline. Если потребуется live e2e — проще
прокинуть свой ToolsetContribution-список из main/test.
Финальный swap — ChatAgent/ChatConversation теперь работают через абстрактный
StorageBundle (pw.binom.agentik.storage), а не через конкретный SqliteStores.
Подготовка к Android-портированию (там будет :storage-android вместо
:storage-sqlite).
Изменения:
- SqliteStores.asBundle() — convenience для превращения конкретного
SQLite-импла в StorageBundle
- ChatAgent(private val storage: StorageBundle) — было stores: SqliteStores
- ChatConversation(private val storage: StorageBundle) — то же
- Main.kt, DebugRoutes.kt — вызовы обновлены, используется .asBundle()
- Все 5 тестовых файлов с ChatAgent(... stores = ...) — обновлены на
ChatAgent(... storage = ...asBundle())
- StorageBundle : AutoCloseable — закрывает все 4 store'а; в тестах
tearDown { storage.close() }
Конфиг не менялся: toolsets остаётся emptyList() по умолчанию (полная
невидимость механики тулсетов для модели). Подключение тулсетов — opt-in
через параметр ChatAgent(toolsets = ...) для будущего e2e-теста в post-implementation.
Tests: 340/340 green. Fatjar 240 MB. Без регрессий.
Добавлен SystemPromptToolsetSection — рендер markdown-секции для system prompt.
Контракт:
- toolsets пустой → null (секция не добавляется, агент не знает о механике)
- иначе → краткое описание концепции + список 'name — description' для
активных и неактивных (одинаковый формат per design contract)
- auto-activation НЕ упоминается в промпте (только в dispatch)
Интеграция в ChatAgent:
- Добавлен параметр toolsets: List<ToolsetContribution> = emptyList()
- При пустом списке — enable_toolset/disable_toolset НЕ регистрируются,
секция в system prompt НЕ появляется (полная невидимость per A1-α)
- При непустом — тулы регистрируются, секция добавляется
- ToolsetRegistry + ToolsetDispatchPolicy создаются per-agent (один реестр
на все диалоги — состояние 'активные тулсеты' общее)
Интеграция в ChatConversation:
- Новый параметр toolsetDispatch: ToolsetDispatchPolicy? = null
- runToolAndPersist: если задан — вызов идёт через policy (auto-activate
неактивных тулсетов, fallback в base dispatcher для плоских тулов)
- Иначе — старое поведение через toolsByName
Тесты:
- 7 новых в :agent-toolsets (SystemPromptToolsetSection): пустые списки,
только активные, только неактивные, оба, проверка отсутствия auto-activation
упоминания, registry-based рендер, пустой реестр
- 5 новых в :standalone (ChatAgentToolsetsTest): default (пустой) — нет
тулов и секции; non-empty — тулы и секция есть; enable_toolset активирует;
вызов тула из неактивного тулсета — auto-activate; disable_toolset
снимает из active set (но auto-activate на следующем вызове — by design)
Tests: 340/340 green (335 ранее + 5 новых ChatAgent integration)
Перенесён SQLDelight (4 .sq файла, конфигурация databases { AgentikDatabase })
и 5 SQLite-импл классов (SqliteStores, SqliteConversationStore, SqliteMessageStore,
SqliteWorkingMemoryStore, SqliteReflectionStore) из :standalone в новый JVM-only
модуль :storage-sqlite под пакетом pw.binom.agentik.storage.sqlite.
Изменения:
- Новый :storage-sqlite модуль с sqldelight-плагином + sqlite JDBC driver
- Все .sq файлы и Kotlin-классы переехали с переименованием пакета
- ReflectionStore.kt в :standalone (только SQLite-импл) удалён — функционал
живёт в :storage-sqlite/SqliteReflectionStore.kt
- :standalone/build.gradle.kts: убран sqldelight-плагин и конфигурация,
добавлена зависимость :storage-sqlite
- Все импорты в :standalone (8 main + 7 test) перенаправлены на новый пакет
- ReflectionStoreTest.kt переехал в :storage-sqlite/jvmTest (тестирует
internal fun encode/decodeStringArray в :storage-sqlite)
Совместимость:
- SqliteStores доступен по новому пути pw.binom.agentik.storage.sqlite.SqliteStores
- Старые импорты в тестах обновлены (минимум diff — 1 строка на файл)
- В commit 6 ChatAgent переключится на StorageBundle API; SqliteStores
станет деталью реализации :standalone
Тесты: 299/299 green. Fatjar standalone-all.jar 240 MB.
Преимущества:
- :standalone больше не зависит от SQLDelight плагина (легче поддерживать)
- :storage-sqlite может быть заменён/расширен (например, :storage-android)
- Тесты storage-слоя сгруппированы по модулю реализации
Выносим интерфейсы и data-классы истории диалога (MessageStore / WorkingMemoryStore /
ConversationStore / ReflectionStore + соответствующие sealed-иерархии MessageRecord /
WorkingMemoryEntry / Content / ConversationRecord / Reflection + payload-утилиты) из
:standalone в отдельный KMP-модуль :storage-core (pw.binom.agentik.storage).
Цель — подготовка к Android-портированию и подключению альтернативных реализаций
хранилища без затягивания всей :standalone. Дальше (commit 2/3) — :storage-inmemory
и :storage-sqlite как самостоятельные модули, плюс :storage-android (deferred).
Изменения:
- Новый :storage-core (KMP, commonMain only, jvm + native таргеты) — 12 файлов
- StorageBundle агрегатор (conversationStore + messageStore + workingMemoryStore +
reflectionStore; SkillStore живёт в :skills и подключается отдельно)
- 11 файлов импортов в :standalone переключены на новый пакет
- SqliteReflectionStore оставлен в :standalone до commit 3 (зависит от
SQLDelight AgentikDatabase, которую ещё не отвязали от :standalone)
- 4 теста перенесены в :standalone/.../storage/ с обновлённым пакетом
- PayloadTest переехал в :storage-core/commonTest (тестирует чистые типы)
Tests: 264/264 green (179 :standalone + 6 :storage-core + прочие JVM-модули)
litert-api 7 -> 8. addToolResult(callId, name, result: Unit) ->
LiteDelta (несёт текст пост-тул ответа модели + возможные вложенные
tool-calls). caffeine не публикует parent-аггрегатор, поэтому алиасы
в libs.versions.toml указывают на -jvm flavor напрямую.
ChatConversation.runTurn: убран хак currentParts=[Text(' ')] —
вместо него runToolAndPersist(call) -> (callId, resultText) ->
addToolResult() возвращает LiteDelta, цикл идёт по delta.toolCalls.
Никакого 'призрачного' ответа модели в KV-cache после каждого тула.
Live smoke-test (gemma-4-E2B + SigLIP vector backend):
- 'Запомни: работаю на macOS' -> 'Я сохранил информацию о том, что вы
работаете на macOS' (раньше: 'Чем я могу помочь?')
- 'На чём работаю?' -> 'Вы работаете на macOS'
- цепочка имя->а necdoт -> модель осмысленно продолжает, не сбрасывается
- тесты: 264/264 зелёных
:standalone декларировал зависимость a2a-server, но mount не было (e2e
нашёл 404 на /a2a). A2aBridge (AgentHandler) гоняет A2A-context на
:proto-диалог: contextId -> Conversation (пустой/неизвестный -> новый),
ответ = склеенные AppendText хода (подписка на events() до send, стоп по
End/Interrupted/Error), id диалога в metadata.agentikConversationId.
kaml encodeToString не ставит завершающий перевод строки, из-за чего
сериализованный SKILL.md выглядел так:
---
name: x
description: "y"---
и SkillParser.parse находил MissingClosingFence — каждый скил,
сохранённый через skill_save / SkillMiner, становился нечитаемым после
рестарта агента (каталог терял скил).
Поставлен явный '\n' перед закрывающим fence. Добавлены юнит-тесты
round-trip serialize->parse (обычный, пустой body, спецсимволы YAML).
SkillMiner (сетка безопасности skill self-improvement): каждые
AGENTIK_SKILL_MINING_INTERVAL user-ходов (default 15) LLM смотрит
последние AGENTIK_SKILL_MINING_MAX_TURNS ходы (default 30) + каталог
существующих скилов и возвращает structured JSON {"skills":[...]}.
Найденное upsert-ится в SkillStore — модель "забыла" вызвать
skill_save в ходе разговора, минер добирает её постфактум.
- SkillMiner.kt: короткий LiteConversation (one-shot), blocking-инференс
на Dispatchers.IO, defensive парсинг (кривой ответ -> пустой список).
- SkillMiningPrompts/SkillMiningParser: тот же подход, что
ReflectionParser (structured-output вместо tool-calling).
- ChatConversation.scheduleSkillMining() — хук после каждого хода
(рядом со scheduleReflection); ChatAgent/Main — прокидывание.
- DebugRoutes.kt: AGENTIK_DEBUG_ENDPOINTS=1 включает POST
/debug/reflect, /debug/skill-mine, /debug/curate, /debug/compact и
GET /debug/tokens для ручного триггерирования фоновых фич.
- ChatConversation.forceCompactNow(): принудительный compaction
без проверки порога (для /debug/compact).
- Тесты: SkillMinerTest (5) + SkillMiningParserTest (9); FakeLiteLlm
теперь записывает send()/sendContents() в lastContents.
179 jvm-тестов :standalone зелёные, README обновлён.
LiteLlm API не отдаёт split prompt/completion наружу через send()
(внутренний OpenAI Usage сидит в pw.binom.litert.openai и недоступен),
поэтому измеряем через LiteConversation.tokenCount():
input = tokenCount() до первого send в turn'е
(= system + вся история + tools + только что добавленное user-сообщение)
output = tokenCount() после завершения turn'а - input
(= assistant text + tool calls + tool results за tool loop)
Пишем в assistant-запись как TurnTokens(input, output) в payload_json.
Никаких schema-миграций: payload-формат уже обёрнут в MessageBodyPayload,
просто добавлено опциональное поле tokens.
MessageStore.tokenStats(conversationId) → TokenStats(turns, inputTokens, outputTokens).
На старте агент печатает сводку по всем диалогам:
tokens: 17 convs, 134 turns, in=523844, out=58290, total=582134
Бэкенды без tokenCount() (off-line LiteRT-LM модели) → tokens=null,
старые assistant-записи без метрики → пропускаются в tokenStats без ошибок.
Tests: TokenStatsTest (5 green) + PersistenceTest (unchanged) → 165 total.
Backward compat: legacy plain-array payload всё ещё читается, tokens=null.
Тулзы были написаны ранее (SkillSaveTool/SkillDeleteTool, SkillToolsFactory),
но не были подключены к ChatAgent. Этот коммит закрывает пробел:
* ChatAgent: новый параметр skillStore: SkillStore? = null. Когда задан —
в allTools добавляются SkillToolsFactory.create(skillStore) → агенту
доступны skill_save и skill_delete (помимо read_skill который всегда
есть при непустом каталоге).
* Main.kt: если config.skillsDir задан — создаём DiskSkillStore(File(dir))
и скармливаем агенту. Каталог используется и для чтения (SkillCatalog),
и для записи (DiskSkillStore.upsert/remove) — одни и те же файлы,
никаких рассинхронов между read_skill и skill_save.
* SkillLoader.loadDirectory больше не нужен в Main — DiskSkillStore сам
подгружает каталог в init. Удалён старый импорт.
* Тесты SkillToolsTest (5): SkillSaveTool persists file and surfaces in
catalog; rejects blank name; SkillDeleteTool archives (rename to
.archived); errors on missing skill; colon-named skills map to nested
dirs (backend:spring:db-base → backend/spring/db-base/SKILL.md).
* README: раздел "Навыки" расширен описанием трёх тулов (read_skill /
skill_save / skill_delete).
Smoke: standalone запускается с пустым AGENTIK_SKILLS_DIR, видит
"skills: 0 loaded from /tmp/skills-smoke" (DiskSkillStore создаёт каталог
при отсутствии). Агент при наличии skillStore имеет в своём распоряжении
все три тула для self-improvement'а.
Теперь Phase 3 (skill self-improvement) реально работает end-to-end:
агент может дёрнуть skill_save когда понимает что задача повторяется,
потом в следующих диалогах использовать новый скил через read_skill.
Tests: 160 standalone JVM (+5), 241 всего JVM, 73 native, all green.
Self-reflection: каждые N пользовательских ходов агент запускает one-shot
LLM-размышление о качестве своих ответов, сохраняет score+weakSpots в SQLite,
подмешивает top-K последних рефлексий в system prompt как "слабые места".
* sqldelight: новая таблица reflection (id, conversation_id?, created_at,
turns_analyzed, score, summary, weak_spots_json) + индексы по created_at и
conversation_id.
* persistence: Reflection data class + ReflectionStore interface +
SqliteReflectionStore (insert/get/listRecent/listForConversation/
deleteOlderThan/count + events Flow). weakSpots хранятся как JSON-массив,
парсятся ручным сканером (без kotlinx-serialization в этом модуле).
* agent: LlmReflector (one-shot LiteLlm через createConversation +
send, structured-output JSON). ReflectionParser (hand-rolled,
толерантный к ```json fences и лидирующему/завершающему тексту;
score принимает int или строку; weakSpots — массив).
* agent: ReflectionPrompts (Russian system+user prompts, аналогично
LlmMemoryReviewer/ReviewPrompts).
* agent: ChatAgent.buildSystemPrompt расширен параметром reflections —
добавляется секция `## Self-reflection: твои слабые места за последнее время`
после memory и перед soul-prepend.
* agent: ChatConversation.scheduleReflection — каждые reflectionInterval
пользовательских ходов (счётчик через workingMemory.list) запускает
reflector на Dispatchers.IO, результат сохраняет в reflectionStore
с conversationId. Не блокирует turn.
* AgentikConfig: новые поля reflectionInterval (env AGENTIK_REFLECTION_INTERVAL,
default 10, clamped 0..1000) и reflectionTopK (env AGENTIK_REFLECTION_TOP_K,
default 3, clamped 0..20).
* Main.kt: если reflectionInterval > 0 — создаём LlmReflector(llm); загружаем
top-K из SQLite в system prompt.
* Ids.reflection() — генератор id "refl-<uuid>".
* Tests: ReflectionParserTest (7: clean JSON, fences, лидирующий текст,
score-as-string, отсутствие score, невалидный JSON, escape-последовательности),
ReflectionStoreTest (7: round-trip, listRecent с лимитом, фильтр по
conversation_id, deleteOlderThan, count, encode/decode строк), и
ChatAgentReflectionTest (3: секция скрыта при пустых, присутствует с
score+spots, порядок soul→memory→reflection).
* README: новые env-переменные, раздел "Self-reflection", startup output.
Smoke test подтверждает: reflection секция появляется в system prompt когда
в SQLite есть записи (listRecent возвращает непустой список). При
reflectionInterval=0 reflector не создаётся, scheduleReflection — no-op.
Tests: 305 total green (+17: 7+7+3). Fatjar собирается, logback-вывод
работает (logging: см. предыдущий коммит f5a551b).