Skip to content

MLflow 3.15 GenAI evaluation for image generation + editable criteria - #2

Open
jacksandom wants to merge 14 commits into
mainfrom
feat/mlflow-eval
Open

MLflow 3.15 GenAI evaluation for image generation + editable criteria#2
jacksandom wants to merge 14 commits into
mainfrom
feat/mlflow-eval

Conversation

@jacksandom

@jacksandom jacksandom commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator

What & why

Pivots image evaluation onto MLflow 3.15 GenAI judges and makes the criteria user-manageable. Previously scoring was a single un-calibrated structured model call with no experiment view, history, or control. Now every generated image is scored by multimodal make_judge {{ trace }} judges (which see the image via the get_span_image tool), traced to an MLflow experiment, surfaced in the gallery, with the criteria editable from Settings.

Highlights

  • MLflow eval engine — 5 metric dimensions (Quality, Prompt [merged prompt+purpose], Text, Safe, + a user-defined Custom), each a {{ trace }} judge on the configured multimodal model (reasoning_effort="none", required for tool-calling). Judge scores mirror back into the existing generated_images columns so the gallery UX is unchanged; enrichment (description/tags) stays a structured call.
  • Per-metric hover rationales — each score cell shows the judge's reasoning on hover (metrics._rationales).
  • Consistent issues/improvements — dropped the redundant "Issues found" list; a real Improved prompt is generated from weak-dimension judge feedback whenever a score is weak.
  • Editable eval criteria (Settings) — new eval_criteria table + /api/eval-criteria CRUD; built-in instructions editable (with reset-to-default), Custom fully user-defined (label + instructions + 0-5), disabled by default. Edits apply without an app restart.

Notable fixes

  • Databricks Apps egress: MLflow trace attachments default to a storage-proxy host serverless/Apps can't reach; fixed via experiment artifact_location on a UC Volume (post-deploy) — no egress change.
  • apx toolchain: app now builds on the current apx (app-module, local dist_dir) + added scripts/build-frontend.sh (Vite build incl. Tailwind).
  • Schema bootstrap hardening: _bootstrap_schema runs statements individually, so an owner-only failure (CREATE EXTENSION) no longer skips later DDL — new tables self-create on existing databases.

Notes for reviewers

  • Deploy uses a per-workspace bundle target with coordinates supplied at deploy time; the committed databricks.yml keeps public placeholders (SP-only auth model unchanged).
  • The app currently runs a manually-built wheel because apx build's orval codegen step is broken with the installed apx version (pre-existing, unrelated); scripts/build-frontend.sh documents the UI build in the meantime.

This pull request and its description were written by Isaac.

Replace the app's image scoring with MLflow tracing + make_judge {{ trace }}
multimodal judges, while preserving every gallery field. The structured
ImageAnalyzer call is kept for enrichment (description, tags, missing
elements, improved_prompt, criteria_evaluation) — traced so the image is
captured as an mlflow-attachment:// reference — and the calibratable scores
(0-5 metrics + safety/brand flags) now come from judges that see the image
via get_span_image. Judge outputs are mirrored back into the existing
generated_images columns, so the gallery UI is unchanged.

- new backend/core/_eval.py: setup_mlflow, build_judges (gpt-5-5 +
  reasoning_effort=none), ImageEvaluator (downscale -> traced analyze ->
  get_trace -> judges -> log_feedback -> merge). Image never enters judge
  inputs (context-overflow lesson from the spike).
- startup wiring in _image_gen_dep.py: non-fatal; app.state.image_evaluator
  is None if eval disabled/unavailable -> falls back to plain analyzer.
- per-generation (generate._save_and_analyze) and backfill/import
  (gallery) prefer the evaluator, fall back to analyzer on any failure.
- new POST /gallery/{batch_id}/evaluate: on-demand batch eval under one
  MLflow run for comparison/history.
- config: eval_experiment/eval_enabled/judge_model/judge_reasoning_effort/
  eval_image_max_px (DATABRICKS_VISION_*); deploy wiring in app.yml.template,
  render-app-yml.sh, databricks.yml var, post-deploy.sh (create experiment +
  grant app SP CAN_EDIT). mlflow[databricks]>=3.15.0 added to app deps.
- docs/SPIKE-mlflow-eval.md carried over as design rationale.

Pure logic (downscale, score coercion, merge, judge-input safety) unit-
tested locally. Live 3.15 paths (async autolog attachment, get_span_image
from the app process, log_feedback) to be confirmed on deploy.

Co-authored-by: Isaac
The committed app targeted an older apx: pyproject used the app-entrypoint
metadata key and _factory/_static imported dist_dir from the apx-generated
_metadata.py. Current apx versions require the app-module key and no longer
emit dist_dir, so `apx build` failed (app-module missing, then dist_dir
ImportError) — the app was un-buildable on the current toolchain,
independent of the MLflow work.

- rename [tool.apx.metadata] app-entrypoint -> app-module (nothing imports
  the symbol; apx uses it internally).
- derive dist_dir in _config.py (same resources.files(app_slug)/"__dist__"
  path static_assets_path already uses) and import it there from _factory
  and _static, instead of from the apx-generated _metadata.py.

Verified create_app() imports cleanly with the new-apx _metadata shape.

Co-authored-by: Isaac
Live deploy showed judges failing with "Must specify 'trace'": get_trace
was called immediately after the span closed, but trace export is async,
so it returned an unflushed/corrupted trace. Use
get_trace(trace_id, flush=True) to force pending writes to complete first.

Necessary but not alone sufficient in the Databricks Apps runtime, where
outbound egress to the MLflow artifact-storage host is refused (see the
follow-up on trace storage).

Co-authored-by: Isaac
Research (JIRA FEVM-391 + the FE Slack thread) identified this exact
issue in FEVM workspaces and its verified fix. MLflow uploads trace
attachments (the image get_span_image fetches) to the experiment's
artifact store; the default is the regional blob proxy
(*.storage.cloud.databricks.com), which serverless/Apps compute cannot
reach (internet-zone signed URL -> Connection refused). Opening egress
is a no-op — the proxy is unreachable by design.

Fix, all within workspace-admin scope: create the eval experiment with
artifact_location on a UC Volume (a path the app already reaches), and
grant the app SP READ/WRITE on it. post-deploy.sh now creates the
Volume + experiment + grants; setup_mlflow just binds to that
experiment (dropped the UC-table trace_location / SQL-warehouse
approach, which fixed span export but not attachments).

Also: post-deploy steps 2-4 now run via `uv run --no-project` so they
don't resolve the app's mlflow>=3.15 dep (unavailable on some networks).

Co-authored-by: Isaac
MLflow multimodal-judge eval is live in the app: visual_quality and the
text judges score real generations, with results mirrored to the gallery.
Document the Databricks Apps egress finding (FEVM-391) and the fix —
create the experiment with artifact_location on a UC Volume so trace
attachments are reachable — plus the wheel-caching and
contaminated-experiment gotchas. Clarify the eval_experiment bundle var.

Co-authored-by: Isaac
purpose_fit and text_legibility were still coming from the structured
ImageAnalyzer call; only 3 of the 5 metric dimensions were judge-scored
(so a no-criteria generation showed 3 assessments on the trace). Add
purpose_fit and text_legibility as always-on {{ trace }} judges and map
their scores onto the metrics dict in _merge, so all five gallery metrics
are judge-derived. brand_conflict + criteria_adherence remain conditional
on criteria being supplied.

Enrichment (description, tags, missing_elements, improved_prompt) stays a
structured call — make_judge emits a scalar Feedback and can't produce
those. Cost: 5 always-on judge calls per generation (was 3), still off the
hot path and parallelised.

Co-authored-by: Isaac
Each of the five metric judges returns a per-dimension rationale, but only
prompt_adherence's was kept (as the evaluation paragraph); the other four
were discarded. Collect all five in _merge (keyed by metric key,
visual_quality -> quality) and nest them under metrics._rationales via a
shared metrics_json() helper used at every DB write site. The image dialog
now shows each dimension's rationale as a hover tooltip on its metric cell
(native title=, matching the app's existing tooltip style), with cursor-help
when a rationale is present. Degrades cleanly to "label: score/5" on the
structured-analyzer fallback (no _rationales key).

No schema/model/API change — metrics is already dict|None so the nested key
passes through.

Co-authored-by: Isaac
The previous commit's __dist__ was built without the @tailwindcss/vite
plugin, so all utility classes were dropped and the app rendered unstyled.
Add a committed frontend-build script that includes BOTH @vitejs/plugin-react
and @tailwindcss/vite (mirroring app/.apx/plugin.ts), so rebuilding the UI
without the broken apx build orval step produces correctly-styled assets.
Documents the Tailwind requirement so the CSS-drop regression can't recur.

Co-authored-by: Isaac
The five metric scores (each with a per-dimension judge rationale on hover)
already convey what's weak, so there's no separate "Issues found" list.

"Improved prompt" is a genuine rewrite, not a restatement of the eval: when
judges flag any dimension below threshold and the structured analyzer
returned no rewrite, generate a real rewritten generation prompt grounded in
the weak dimensions + their rationales (analyzer's own client/model, a
background text call on weak images only). This honors "anything with issues
gets an improved prompt".

Safety/brand flags — not captured by the numeric scores — show as a small
inline warning only when present. missing_elements is still stored but no
longer displayed (redundant with scores).

Co-authored-by: Isaac
Replace the hardcoded 5 metric judges with a DB-managed fixed set of 5
criteria (eval_criteria table, seeded via _SCHEMA_SQL). Four built-ins
(Quality, Prompt [merges the old prompt-adherence + purpose-fit], Text,
Safe) have editable instructions with reset-to-default; the 5th is a
user-defined Custom criterion (label + instructions + 0-5), disabled by
default until enabled.

- _eval.py: DEFAULT_CRITERIA + load_criteria(pool); build_judges builds one
  {{ trace }} judge per enabled criterion; ImageEvaluator loads criteria
  from the DB and exposes refresh_criteria(pool) so Settings edits apply
  without a restart. _merge maps built-ins to their metric keys and nests
  the custom result under metrics._custom; purpose_fit is no longer written.
- New routers/eval_criteria.py: GET list, PATCH {key} (instructions;
  custom also label/enabled), POST {key}/reset — each refreshes the live
  evaluator. models.py: EvalCriterionOut/Update. Registered in router.py.
- Frontend: Settings "Eval Criteria" tab (edit 4 built-ins + reset, edit +
  enable Custom); hand-added eval-criteria hooks/types to lib/api.ts
  (orval build step is broken); image-dialog grid = 4 fixed cells + a
  dynamic Custom cell from metrics._custom, each with its hover rationale.

Co-authored-by: Isaac
Two live-deploy bugs in the eval_criteria feature:
- load_criteria indexed rows positionally (r[0]...), but the app pool uses
  dict_row, so it raised KeyError: 0 -> "load_criteria failed (0)" and fell
  back to DEFAULT_CRITERIA (custom disabled), so edits/enable never applied.
  Force dict_row + access by column name.
- The app runs uvicorn --workers 2; a Settings PATCH only refreshed the
  worker that served it, leaving the other worker's evaluator stale. Reload
  criteria at the start of each (background) score_image and rebuild judges
  only when the criteria signature changed, so every worker self-heals.

Co-authored-by: Isaac
_bootstrap_schema ran the whole _SCHEMA_SQL as one multi-statement query,
so a single failing statement (e.g. CREATE EXTENSION -> "must be owner" for
a non-owner service principal) aborted the implicit transaction and skipped
every statement after it — which is why newly-added tables like
eval_criteria weren't created on existing databases (only post-deploy.sh,
run as the owner, created them).

Split _SCHEMA_SQL into individual statements (a scanner that ignores
semicolons inside single-quoted literals — there are 27, and the
eval_criteria seed must stay one statement; no dollar-quoted blocks exist)
and execute each on the autocommit connection with per-statement error
handling. Expected "already exists"/"must be owner" errors are logged at
debug; all DDL is idempotent so this is safe on every startup. New tables
now get created even when earlier owner-only statements fail.

Co-authored-by: Isaac
Remove concrete deployment identifiers (workspace host + id, catalog/schema
name, Lakebase/Postgres paths, app URL, batch job id) from the spike doc and
replace with generic placeholders / a pointer to the bundle variables. The
design rationale is unchanged; nothing workspace-specific remains committed.

Co-authored-by: Isaac
- Remove docs/SPIKE-mlflow-eval.md (spike artifact, not part of the shipped
  solution) and its dangling reference in _eval.py's module docstring.
- README: describe MLflow 3.15 GenAI judge evaluation across the intro,
  What it does, Stack (new Evaluation line + trace-attachment storage note),
  What's interesting, and the dependency table; add MLflow (judges + tracing,
  UC-Volume trace attachments, editable criteria) to the architecture diagram
  and single-image flow.

Co-authored-by: Isaac
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant