Add 2026-06-llmops-quickstart blog code - #91
Conversation
End-to-end LLMOps quickstart on Databricks (customer support ticket classifier): MLflow ChatAgent on a Foundation Model API endpoint, an evaluation gate that promotes a Champion in Unity Catalog, and a Databricks Asset Bundle that deploys schema, experiment, and jobs with batch + real-time inference. - Folder follows the YYYY-MM-[name] convention - README documents setup, structure, data (30 synthetic tickets, no PII), and licenses - LICENSE.md is an unmodified copy of the repo Databricks license - CODEOWNERS entry added for the new folder I have read the contribution guidelines. No sensitive info, no PII, no external dataset. Pending: SME code review approval and internal approval. Co-authored-by: Isaac
Co-authored-by: Isaac
…al, AI Gateway logging Bring the quickstart current (2026) and align it with the MLOps quickstart's step-by-step, best-practices structure. Tested end-to-end across three workspaces (deploy → 4 jobs → approval → governed endpoint → batch + realtime). - Evaluation: replace the hand-rolled accuracy loop with mlflow.genai.evaluate() using a deterministic exact_match gate scorer plus the built-in Correctness LLM judge; every row is captured as an MLflow Trace. - Lifecycle: passing versions register as Challenger; new model_approval.py promotes Challenger→Champion only when run with --params approved=true, wired as a predecessor task to deployment. - Governance: after agents.deploy(), enable AI Gateway inference-table payload logging on the agent endpoint (idempotent). Guardrails/rate limits documented as an FM-endpoint pattern (not supported on custom agent endpoints). - Agent: normalize response content so reasoning models (Claude Sonnet 5, GPT-5) that return structured content blocks work, not just plain-string responses. - Default LLM endpoint → databricks-claude-sonnet-5; pin mlflow>=3.4.0 in the logged model. - Harden stale-deployment cleanup; README updated with the new flow, approval gate, governance notes, UC-privilege prerequisite, and a "before you call it done" checklist.
Per SME review, move off Model Serving / agents.deploy() to the current recommended pattern: serve the agent as a Databricks App and call the LLM through a Unity AI Gateway model service. Validated locally (agent server serves correctly) and against real model services on two workspaces; MLflow 3 GenAI eval passes at 90%. - agent_server/: FastAPI app (mlflow.genai.agent_server @invoke handler) classifying a ticket into one of five categories; calls the model service by fully-qualified name via the AI Gateway; reasoning-model-safe content handling. - LLM via UAIG model service (LLM_MODEL = <cat>.default.claude-sonnet-5; also gpt-oss-120b); governance lives in Unity Catalog. - Evaluation gates the app deploy: mlflow.genai.evaluate with an exact_match gate scorer + the out-of-the-box Correctness judge; agent-evaluate exits non-zero below threshold. Dropped UC model registration and Champion/Challenger aliases. - Bundle: app resource + experiment + data-ingestion job (writes the support_tickets UC managed table). uv.lock via --exclude-newer for proxy-mirrored versions. - Current naming throughout: Declarative Automation Bundles, UAIG, UC managed table, out-of-the-box. README adds prerequisite-knowledge, qs_catalog naming, training pointers, UC-privilege note, and a "before you call it done" checklist.
Link the free DevOps Essentials for Data Engineering (bundles/CI-CD) and Building Agentic Applications on Databricks (agents, MLflow, evaluation) courses, per SME suggestion to point readers at relevant training.
Co-authored-by: Isaac
kwulffert23
left a comment
There was a problem hiding this comment.
Hi @CEDipEngineering , I left a couple of comments. Please let me know if you have questions.
There was a problem hiding this comment.
@CEDipEngineering, not sure if you want to keep or rework the diagram to align with the code, as it's showing the agent deploy as model instead over apps.
There was a problem hiding this comment.
Good catch, updated the diagram.
There was a problem hiding this comment.
@CEDipEngineering, same as above, I'd suggest reworking the diagram to align with the code, as it's showing the agent as mlflow model instead of agent on apps.
There was a problem hiding this comment.
Both diagrams don't seem to be mentioned in the readme...
There was a problem hiding this comment.
Updated both diagrams, and added them to readme.
There was a problem hiding this comment.
@CEDipEngineering the current .env.example is missing CATALOG_NAME and SCHEMA_NAME and LLM_MODEL
There was a problem hiding this comment.
Right, forgot to update the .env, should be fixed now.
|
|
||
| ```bash | ||
| databricks bundle deploy \ | ||
| -v catalog_name=qs_catalog \ |
There was a problem hiding this comment.
@CEDipEngineering I might be outdated with the doc but isn't it --var?
There was a problem hiding this comment.
Yes, it is, hehe. Fixed here and on the Blog.
| w = WorkspaceClient() | ||
| resp = w.api_client.do( | ||
| "POST", | ||
| f"/api/2.0/apps/{app_url}/invocations", |
There was a problem hiding this comment.
@CEDipEngineering could you double check the invocation code, did it work? I had in mind /api/2.0/apps/{app_name}/invocations from https://github.com/databricks/databricks-sdk-py/blob/main/databricks/sdk/service/apps.py (line 3398)
There was a problem hiding this comment.
/api/2.0/apps/{app_name}/invocations does not work. I mixed it up with {app_url}/invocations, which is the endpoint I built on the app.
| resources: | ||
| apps: | ||
| ticket_classifier: | ||
| name: "llmops-quickstart-classifier" |
There was a problem hiding this comment.
do we need a prefix in the app's name? e.g. if we deploy in the same workspace without any prefix the same name will be hardcoded in both targets dev and prod ending in resource already exists if running on the same workspace. It could be something like name: "llmops-quickstart-classifier-${bundle.target}"
| variables: | ||
| environment: dev | ||
|
|
||
| prod: |
There was a problem hiding this comment.
@CEDipEngineering Shall we add root_path? Setting root_path pins prod to a single location regardless of who runs the deploy. https://github.com/databricks/devhub/blob/main/.agents/skills/databricks-core/declarative-automation-bundles.md
There was a problem hiding this comment.
I added it, however I decided to stick with the same approach as the MLOps repo I'm basing this on, which does root_path = /Workspace/Users/${workspace.current_user.userName}/.bundle/${bundle.name}/${bundle.target}
by default. I also added a comment there to explain why users might want to change this to something static, like /Workspace/Bundles/<project_name>/<bundle_name>/ and why using /Workspace/Shared is not a good idea.
… naming, prod root_path, diagrams
- Replace invalid `-v key=value` with `--var="key=value"` (the -v shorthand
does not exist; the documented commands failed outright)
- Fix the inference snippet: a Databricks App is called on its own URL, not via
a /api/2.0/apps/.../invocations control-plane endpoint (no such API)
- Add the missing LLM_MODEL, CATALOG_NAME, SCHEMA_NAME (plus optional
ACCURACY_THRESHOLD, DATABRICKS_WAREHOUSE_ID) to .env.example
- Suffix the app name with ${bundle.target} so dev and prod can coexist in one
workspace (development mode does not prefix app names)
- Pin the prod bundle root_path to a fixed identity instead of the deployer's
home folder, keeping it out of world-writable /Workspace/Shared
- Regenerate both diagrams to match the App + UAIG model service architecture
and reference them from the README; add the Mermaid sources
- Read the app source path from `bundle summary` rather than hardcoding it
Co-authored-by: Isaac
…he code - app.yaml set LLM_MODEL to a bare endpoint name rather than a model-service fully-qualified name, contradicting the rest of the quickstart; note that the bundle's config.env overrides app.yaml when deploying through the bundle - evaluate_agent.py required no CATALOG_NAME/SCHEMA_NAME and silently fell back to main.llmops_quickstart, so evaluating a prod deploy would score the dev schema and report a green gate against the wrong data - evaluate_agent.py now fails with a clear message when MLFLOW_EXPERIMENT_ID is unset, rather than a RESOURCE_DOES_NOT_EXIST stack trace, and resolves .env relative to the package instead of the current directory - agent.py docstring claimed an OpenAI client is used; it uses the SDK's HTTP client - databricks.yml: prod_service_principal needs a default, since a variable with no default is required by every target and broke dev validation Verified: dev and prod targets validate, and the evaluation runs end to end at 90.0% exact-match against the 30-ticket set. Co-authored-by: Isaac
Sets root_path explicitly on dev and prod with ${workspace.current_user.userName},
the same pattern the companion MLOps Quickstart uses, so the two posts stay
consistent. Drops the prod_service_principal variable, so both targets validate
with no flags.
Leaves a comment recommending a static folder with controlled access, recreated
the same way on every workspace (e.g. /Workspace/Bundles/<bundle>/<target>), as
the stronger option for a real prod deployment.
Co-authored-by: Isaac
What this is
End-to-end LLMOps quickstart on Databricks — a customer support ticket classifier that carries an LLM agent through the full lifecycle: data ingestion → agent build → evaluation → approval → deployment → inference.
Accompanies an upcoming Databricks Community blog post (JIRA TLC-1077).
Architecture (2026 pattern)
mlflow.genai.agent_server@invokehandler), not a Model Serving endpoint. Serving agents as apps is the current recommendation;agents.deploy()is reserved for special custom cases.<catalog>.default.claude-sonnet-5, alsogpt-oss-120b) through the AI Gateway. The model service is a Unity Catalog securable, so access control, rate limits, and payload logging live in UC. One env var (LLM_MODEL) switches models.mlflow.genai.evaluatewith a deterministicexact_matchgate scorer plus the out-of-the-boxCorrectnessjudge; every prediction is captured as an MLflow Trace.agent-evaluateexits non-zero below the threshold (CI gate). A human approves before the app is deployed.Contents
2026-06-llmops-quickstart/(YYYY-MM-[name]convention)agent_server/(agent, evaluation, server),app.yaml,pyproject.toml+uv.lockdatabricks.yml) with the app resource, MLflow experiment, and the data-ingestion job (dev/prod targets)README.mdwith prerequisite-knowledge, setup, governance recap, and a "before you call it done" checklistLICENSE.md(unchanged) andCODEOWNERSTesting
/invocationsreturns correct categories.mlflow.genai.evaluaterun: 90% exact-match, gate passed, traces logged.Guidelines checklist
YYYY-MM-[folder_name]Blog post link will be added to the README once published.
This pull request and its description were written by Isaac.