Rerun the analysis when new data lands and send the figures that changed
For a standing research question: the dot watches a data folder, reruns the notebook on each new file, flags anomalies and sends only the figures that moved.
Jev judge eval (1 score) over a fixed state : runnable via experimental_evaluate.
Community-reported results. Success rates appear after 10 approved reports. Copies are deduplicated per visitor each day. Starter workflows have not been independently tested.
Jev judge eval (1 score) over a fixed state : runnable via experimental_evaluate.
Copy this into Muse and replace anything in [square brackets] with your own details.
// Grade: project plan completeness — Jev judge eval (typesafe-ai/jev).
// An agent drafts a launch plan; the judge scores its completeness.
// Model: typesafe-ai/jev via the Vercel AI Gateway ($0.042/1M input tokens).
// Auth: set AI_GATEWAY_API_KEY (vck_…) in the environment — never paste the key here.
// For zero-data-retention / no-training workloads, add
// providerOptions: { gateway: { … } } to the evaluate call below.
import { experimental_evaluate } from 'ai';
const result = await experimental_evaluate({
model: 'typesafe-ai/jev',
state: "Plan: 1) Freeze scope Friday. 2) Ana owns staging QA, Ben owns migration dry-run. 3) Rollback: one-click revert to snapshot v41, tested Thursday. 4) Comms: status page + email at T-24h and T+1h. 5) Success: error rate <0.5% and p99 < 400ms for 48h.",
questions: {
"completeness": {
"type": "score",
"instructions": "Score the completeness of this launch plan. Higher is better.",
"criteria": [
"A bare task list with no owners, dates, or rollback.",
"Tasks plus owners, but no rollback, comms, or success criteria.",
"Owners, dates, and rollback, but vague comms or success measure.",
"Owners, dates, tested rollback, comms plan, and numeric success criteria."
]
}
},
});
console.log(JSON.stringify(result, null, 2));
Source / inspiration: View credited source
Help the next person know what to expect.
Sign in to share your resultA question, a useful tweak, a little discovery. Start the conversation.
Sign in to join inFor a standing research question: the dot watches a data folder, reruns the notebook on each new file, flags anomalies and sends only the figures that moved.
An organized catalog of relevant datasets with access, freshness, and licensing notes.
Review what the agent will be allowed to touch: browse the web, read email you connect. Ask Muse to show its plan and pause before anything that spends money, sends a message or changes an account.
Reported by a Muse user on vercel.com (original post). Collected by the open-source EveryAI Muse field guide under the MIT licence. Sundog has not tested this workflow; results people report here appear on this page.