feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5) - #478
Open
MarioCadenas wants to merge 1 commit into
Open
feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5)#478MarioCadenas wants to merge 1 commit into
MarioCadenas wants to merge 1 commit into
Conversation
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
from
August 18, 2026 12:57
0b55719 to
1431ce7
Compare
Contributor
📦 Bundle size reportCompared against
|
| dist | raw | gzip |
|---|---|---|
| JS (runtime) | 1.1 MB (+35 KB) | 397 KB (+15 KB) |
| Type declarations | 406 KB (+20 KB) | 145 KB (+9.3 KB) |
| Source maps | 2.2 MB (+72 KB) | 745 KB (+30 KB) |
| Other | 11 KB | 3.7 KB |
| Total | 3.7 MB (+128 KB) | 1.3 MB (+54 KB) |
Per-entry composition (own code — deps external (as shipped))
| Entry | Initial (gz) | Lazy (gz) | Total (gz) | node_modules (min) | Own code (min) |
|---|---|---|---|---|---|
. |
95 KB (+1 B) | 2.5 KB | 98 KB (+1 B) | external | 313 KB |
./beta |
81 KB (+4.5 KB) | 455 B (-1 B) | 81 KB (+4.5 KB) | external | 242 KB (+12 KB) |
./testing |
17 KB | 0 B | 17 KB | external | 51 KB |
./tsdown |
520 B | 0 B | 520 B | external | 813 B |
./type-generator |
23 KB | 0 B | 23 KB | external | 65 KB |
Chunks:
| Entry | Chunk | Load | Size (gz) |
|---|---|---|---|
. |
index.js |
initial | 91 KB |
. |
utils.js |
initial | 4.0 KB |
. |
remote-tunnel-manager.js |
lazy | 2.5 KB |
./beta |
beta.js |
initial | 65 KB |
./beta |
stream-manager.js |
initial | 5.8 KB |
./beta |
wide-event-emitter.js |
initial | 3.2 KB |
./beta |
databricks.js |
initial | 3.2 KB |
./beta |
configuration.js |
initial | 2.1 KB |
./beta |
service-context.js |
initial | 1.3 KB |
./beta |
client.js |
initial | 434 B |
./beta |
client-options.js |
initial | 219 B |
./beta |
supervisor-api.js |
lazy | 192 B |
./beta |
databricks.js |
lazy | 141 B |
./beta |
index.js |
lazy | 122 B |
./testing |
index.js |
initial | 17 KB |
./tsdown |
index.js |
initial | 520 B |
./type-generator |
index.js |
initial | 23 KB |
@databricks/appkit-ui
npm tarball (packed): 350 KB (-4 B) — gzipped download (dist + bin; excludes release-only docs/NOTICE).
| dist | raw | gzip |
|---|---|---|
| JS (runtime) | 395 KB | 132 KB |
| Type declarations | 229 KB | 84 KB (-1 B) |
| Source maps | 766 KB | 253 KB |
| CSS | 16 KB | 3.2 KB |
| Total | 1.4 MB | 473 KB (-1 B) |
Per-entry composition (consumer bundle — deps bundled, peerDeps external)
| Entry | Initial (gz) | Lazy (gz) | Total (gz) | node_modules (min) | Own code (min) |
|---|---|---|---|---|---|
./js |
5.3 KB | 49 KB | 55 KB | 208 KB | 14 KB |
./js/beta |
20 B | 0 B | 20 B | 0 B | 0 B |
./react |
432 KB | 49 KB | 481 KB | 1.3 MB | 177 KB |
./react/beta |
1.0 KB | 0 B | 1.0 KB | 0 B | 1.9 KB |
Chunks:
| Entry | Chunk | Load | Size (gz) |
|---|---|---|---|
./js |
index.js |
initial | 5.2 KB |
./js |
chunk |
initial | 120 B |
./js |
apache-arrow |
lazy | 49 KB |
./js/beta |
beta.js |
initial | 20 B |
./react |
index.js |
initial | 430 KB |
./react |
tslib |
initial | 2.1 KB |
./react |
apache-arrow |
lazy | 49 KB |
./react/beta |
beta.js |
initial | 1.0 KB |
Contributor
🤖 AppKit PR bot🔬 Run evalsStart an eval for this PR from the evals-monitor app: Go to Evals Monitor → 📦 Try this PR's app templateScaffolds a new app from this PR's SDK build. Run it in any folder (requires the GitHub CLI — gh run download 33763255268 -R databricks/appkit -n appkit-template-0.71.0-pr.97064b1-pr-agent-evals-2-framework-478 -D appkit-pr-478 \
&& unzip -o "appkit-pr-478/appkit-template-0.71.0-pr.97064b1-pr-agent-evals-2-framework-478.zip" -d "appkit-pr-478" \
&& databricks apps init --template "appkit-pr-478"The template pins |
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
2 times, most recently
from
August 21, 2026 10:37
e32afdd to
65de478
Compare
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
2 times, most recently
from
September 1, 2026 13:20
1f15f16 to
274ec9d
Compare
MarioCadenas
changed the base branch from
main
to
fix/agents-mlflow-single-provider
September 2, 2026 14:39
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
from
September 3, 2026 09:39
da02130 to
5fe6166
Compare
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
from
September 3, 2026 11:59
5fe6166 to
74929cc
Compare
Rebuilt on the current fix/agents-mlflow-single-provider base (#545 was force-pushed onto v0.70.0, so the branch's merged copies were stale). The eval framework is self-contained — it imports nothing from the agents plugin — so every #545 file is taken from the base untouched; only the framework's own files plus additive seam edits (beta/connectors/cli exports, autoevals dep) are applied. Framework: discover server/agents/<agent>/evals/*.eval.ts, drive a running app over HTTP+SSE, run deterministic assertions + LLM-as-judge (autoevals), and report results + per-trace assessments to MLflow (classic V3 + UC V4). Includes the mlflow REST connector and the `appkit agent eval` CLI. Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
from
September 3, 2026 13:48
74929cc to
f8102dd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack 2/5 · targets
pr/agent-evals-1-tracing(review after #1).The core eval framework, plus LLM-as-judge and the MLflow REST connector.
defineEval): drive an agent over HTTP against a running app; assert witht.succeeded(),t.calledTool(),t.check(value, matcher)(includes/equals/matches). Gate-by-default,.soft()to demote.genai_evaluate; each turn's trace links viamlflow.sourceRun; per-assertion feedback written via the assessments REST API.t.judge.factuality/closedQA/custom) via autoevals → a Databricks serving endpoint.connectors/mlflow:MlflowClient(host/token, post/postResult, serving URL) +resolveDatabricksAuth/resolveWorkspaceClient(OAuth from a CLI profile — no hand-set PAT). Extracted so both evals and future callers share the REST/auth layer.appkit agent evalCLI.Squashed history note: contains the framework, judge, and connector-extraction commits.