If you’ve been asked to evaluate a student-facing assistant, you’ve probably noticed that the demos are unhelpfully similar. Every one of them answers a sensible question about an assignment deadline with a fluent, well-formatted reply, and every vendor uses roughly the same three words to describe how it works.
The differences that matter don’t show up in that demo. They show up in week three of term, when a student asks something nobody documented, or asks the assistant to do their coursework, or types something at midnight that should have gone to a human being immediately.
So it’s worth having a specific list of questions, and knowing which answers should end the conversation.
Why this isn’t ordinary customer support
Education has constraints that don’t apply when the same technology sits on a shop’s website, and they change what “good” means.
There’s academic integrity, which has no equivalent elsewhere — a tool that helps a customer too much is a good tool; one that helps a student too much is a misconduct case. There’s the fact that your users can’t leave: a frustrated shopper goes to a competitor, a frustrated student is stuck with the system and takes it out on your support desk and their course evaluation. There’s faculty trust, which is slow to earn and fast to lose. And the load is violently seasonal, arriving in the first fortnight of term and the days before major deadlines, which is exactly when nobody has capacity to fix a bad deployment.
The eight questions
- 1. Where do the answers come from? The only acceptable answer is: from your own course content and help material, retrieved at the moment of the question. If it’s answering from general knowledge about how learning platforms usually work, it will confidently describe a submission process you don’t use.
- 2. What does it do when it doesn’t know? Ask for a policy you’ve never published. A good assistant says it doesn’t know and routes the student onward. A bad one invents a plausible late-submission window, and a student will act on it.
- 3. Will it refuse coursework? This is the question faculty will ask first. Type a graded question from a real assignment and see what happens. It should decline and point back to the material, not reason its way to an answer. If it helps, the evaluation is over — no amount of convenience survives an academic misconduct panel.
- 4. Can it hand off to a person, and does that route actually go somewhere? A handoff into an unmonitored inbox at 11pm is worse than no handoff, because it produces the impression of help without any. Ask what the escalation looks like out of hours, and who is on the other end.
- 5. What happens to what students type? Where is it stored, for how long, who can read it, and does any of it train anything? Students disclose more than they intend in a chat box — including personal circumstances behind an extension request — and you’re accountable for that regardless of what the vendor’s terms say.
- 6. Does it work on the systems you actually run? Very few institutions are single-platform. If one faculty is on a different system, or a migration stalled at seventy percent, an assistant that only speaks one platform’s vocabulary solves half your problem and creates a new one.
- 7. Which languages does it handle, in practice? Not the marketing list — test it. A student who writes in their first language when they’re stressed should get a usable answer, not a polite apology in English.
- 8. Who maintains the content, and how does an answer get changed? Every assistant of this kind mirrors the material behind it. If updating an answer requires a vendor ticket rather than editing your own page, the thing will be out of date within a term.
A thirty-minute test that beats any demo
Vendors demo on clean, curated content. Your estate is not clean or curated, and the gap between the two is where deployments fail. So insist on testing against one of your own real courses — ideally a slightly messy one, from a department that has been rewriting its assessment brief all summer.
Then run four questions in this order:
- Something clearly documented — the control. It should get this right and cite something recognisable.
- Something documented in two places that contradict each other. Watch whether it picks one silently or flags the conflict. This tells you what will happen across a real estate.
- Something genuinely undocumented. The only good outcome is an admission and a handoff.
- A graded question, phrased exactly as a student would try it — including the second attempt where they claim they only want a hint.
You’re not testing whether it’s clever. You’re testing whether it knows the edges of what it should do. Three of those four questions are designed to make it fail, and how it fails is the entire finding.
Surveying what’s actually available
The field is more crowded than most evaluations assume, and it varies by platform. On the Moodle side there’s a mix of free community plugins and hosted services with quite different trade-offs around install effort, language coverage, and how deeply they read course content — worth surveying before you shortlist rather than after, and there’s a comparison of the main Moodle options that lays the differences out side by side.
The Canvas picture looks different, partly because more of it runs through the institution’s own pages and course material rather than a plugin directory. The practical shape of an assistant inside a Canvas course is the same idea — it reads from what you’ve published and answers navigation questions in plain language — but the evaluation questions above matter more than the platform badge. If you run both systems, ask each vendor to demonstrate on both rather than assuming it carries over.
The two failure modes worth planning for
The first is scope creep by accident. It arrives as a reasonable request — could it also explain the concept the student is stuck on, since it’s right there? Each step is defensible and the destination is a tool that is quietly tutoring, which is a conversation you do not want to have with an examinations board. Write the boundary down before launch: navigation and logistics on one side, teaching and assessment on the other.
The second is content rot. An assistant answers from what you’ve written, so a syllabus that contradicts a module page produces a confidently wrong answer at scale, and nobody notices until a student acts on it. Whatever you point it at needs a named owner and a review date — which is unglamorous, and is the real work of the project.
The honest limits
It won’t fix a disorganised course. If submission instructions are missing and the late policy lives in an email from August, an assistant reading that material will relay the confusion back faithfully. Setting one up tends to surface every gap you’d half-forgotten about, which is annoying and genuinely useful in equal measure.
It also won’t replace your support team. Students will keep having odd, specific, human problems that need somebody who can look at an individual record and make a judgment call. What it removes is the repetitive layer on top — where do I submit, when is it due, why can’t I see the next section — which is most of the volume and almost none of the value.
And it won’t help much if it goes live in week one of term. Launch it in a quiet period, on one course, with the transcripts read daily by somebody who can fix the content it exposes.
What you’re actually choosing
Underneath the feature comparisons, you’re choosing how a system behaves when it’s out of its depth — because with students, at midnight, in week three, it will be.
An assistant that answers ninety percent of questions well and bluffs the rest is worse than one that answers seventy percent and says “I don’t know, here’s who can help.” That’s not a nuance of procurement. It’s the whole decision, and every question above is a way of getting at it before a student finds out on your behalf.
