Rerun the analysis when new data lands and send the figures that changed
For a standing research question: the dot watches a data folder, reruns the notebook on each new file, flags anomalies and sends only the figures that moved.
Jev judge eval (1 choice + 1 score) over a fixed state : runnable via experimental_evaluate.
Community-reported results. Success rates appear after 10 approved reports. Copies are deduplicated per visitor each day. Starter workflows have not been independently tested.
Jev judge eval (1 choice + 1 score) over a fixed state : runnable via experimental_evaluate.
Copy this into Muse and replace anything in [square brackets] with your own details.
// Review: interview feedback hire signal + evidence — Jev judge eval (typesafe-ai/jev).
// A panel writes feedback on a backend candidate; the judge reads the signal and checks evidence.
// Model: typesafe-ai/jev via the Vercel AI Gateway ($0.042/1M input tokens).
// Auth: set AI_GATEWAY_API_KEY (vck_…) in the environment — never paste the key here.
// For zero-data-retention / no-training workloads, add
// providerOptions: { gateway: { … } } to the evaluate call below.
import { experimental_evaluate } from 'ai';
const result = await experimental_evaluate({
model: 'typesafe-ai/jev',
state: "Feedback: Strong hire. System design was excellent — she drove the sharding discussion, named the hot-partition risk unprompted, and sketched a clean migration. Coding: correct optimal solution in 25 min with clear tests. Concern: none material; ramp-up on our queue infra expected.",
questions: {
"signal": {
"type": "choice",
"instructions": "Classify the overall hire signal of this feedback.",
"criteria": {
"strong_hire": "Enthusiastic yes with concrete standout evidence.",
"hire": "Yes, with solid but not exceptional evidence.",
"no_hire": "No, with concrete gaps or red flags.",
"mixed": "Contradictory or insufficient evidence either way."
}
},
"evidence_quality": {
"type": "score",
"instructions": "Score how evidence-backed this feedback is. Higher is better.",
"criteria": [
"Verdict with no examples.",
"Verdict with vague praise (great, smart).",
"Verdict with one concrete example.",
"Verdict with multiple specific, observed examples plus a named non-concern."
]
}
},
});
console.log(JSON.stringify(result, null, 2));
Source / inspiration: View credited source
Help the next person know what to expect.
Sign in to share your resultA question, a useful tweak, a little discovery. Start the conversation.
Sign in to join inFor a standing research question: the dot watches a data folder, reruns the notebook on each new file, flags anomalies and sends only the figures that moved.
An organized catalog of relevant datasets with access, freshness, and licensing notes.
Review what the agent will be allowed to touch: browse the web. Ask Muse to show its plan and pause before anything that spends money, sends a message or changes an account.
Reported by a Muse user on vercel.com (original post). Collected by the open-source EveryAI Muse field guide under the MIT licence. Sundog has not tested this workflow; results people report here appear on this page.