What Is the Best Way for a School to Buy an AI Tutor?
The safest approach is to treat an AI tutor as instructional technology, not as an independent teacher or automatic learning solution. A school should first define the academic problem, approve a limited pilot, verify learning outcomes with human-led measures, and set enforceable controls for student data, teacher oversight, accessibility, and vendor claims. By September 2026, the relevant procurement question is no longer simply whether an AI tutor can explain a topic, but whether it improves defined outcomes for defined students without creating unacceptable privacy, safety, equity, or contractual risks. Evidence reported by Third Space Learning indicates that its spoken mathematics tutor Skye has been used to provide mathematics tutoring across schools participating in the UK’s National Tutoring Programme. That is meaningful deployment evidence, but it is not the same as proof that every AI tutor will work in every classroom or subject. The Federation of American Scientists recommends prioritizing student safety in K-12 AI procurement through explicit guardrails, while the Manhattan Institute argues that a single standardized approach is unsuitable because children have different developmental needs. The defensible answer is therefore a controlled, evidence-based purchase rather than a districtwide technology rollout.
Also worth reading: How Can Schools and AI Tutorial Platforms Comply With Educational Data Privacy Rules in 2026? · Which AI classroom pilot metrics should schools measure before scaling AI-driven tutorials? · How Should an AI Tutor Pilot Be Designed for Schools in 2026?
A good procurement process also recognizes that educational value and business value are separate. A platform can be inexpensive per licence, easy to use, and popular with students while still producing weak attainment gains or inappropriate interactions. Conversely, a more expensive system may justify its cost if independent evidence shows that teachers can use it effectively and that student learning improves. Price should be considered together with setup time, professional development, devices, accessibility features, monitoring, data processing, support, and the cost of replacing a product that fails a pilot. No single figure should determine the decision. The chosen product should have a clear instructional purpose, measurable boundaries, and an exit plan.
Which Student Safety and Privacy Rules Should Buyers Require?
Schools should make privacy and safety contractual conditions rather than features that are merely advertised in a sales presentation. At a minimum, a district should ask what student data is collected, why each data field is needed, where information is stored, which companies process it, how long it is retained, and whether the provider uses conversations or student work to train commercial models. The agreement should prohibit the sale of student information and should limit secondary uses unless the district gives specific, informed authorization. It should also define access rights for district staff, vendors, contractors, and any external artificial intelligence services. Families should receive plain-language information about the system rather than a technical policy that is difficult to interpret.
Because AI systems can generate uncertain or harmful responses, schools also need rules for escalation. The tutor should disclose that it is an automated system when users might otherwise believe it is a person, and it should direct students to a teacher or trusted adult for bullying, abuse, medical concerns, crisis disclosures, or requests involving sensitive personal information. Child-development expectations matter: younger pupils may need shorter interactions, more restrictive topics, simpler language, and stronger adult visibility. The Manhattan Institute’s developmentally appropriate framework is relevant here because primary students should not be treated as older users merely because the same interface can technically accept them. Access logs, reporting tools, model-update controls, and prompt filtering should be documented, tested, and available to authorized school personnel.
Cyberinsurance, breach-notification duties, deletion certificates, and security audit rights should be explicit. A district should not assume that a vendor’s public privacy policy is sufficient for FERPA compliance, state student-privacy legislation, or applicable international rules. Legal review is necessary because the tool may involve speech, images, analytics, third-party models, or automated recommendations. The cost of a background check cannot be compared simply with software fees, since a serious data failure could affect students, staff, reputation, and district liability at the same time.
How Should a School Test Learning Effectiveness Before Buying?
A procurement team should begin with a specific instructional problem, such as improving foundational numeracy, providing homework feedback, or offering additional practice in a high-need subject. Broad goals such as “personalized learning” are too vague to test. The school should identify a baseline, select outcome measures, define the comparison method, and decide what result would justify expansion. Relevant measures may include growth on a validated assessment, teacher-observed mastery, completion of supervised practice, attendance, and student well-being. Raw answer accuracy alone is inadequate because a student could guess, copy an answer, or rely on the system without understanding the reasoning.
A controlled pilot should use a sample large enough to reveal meaningful variation but small enough to limit cost and disruption. A practical minimum is one term, although eight to twelve weeks often provides enough time for onboarding, baseline measurement, instructional use, assessment, and analysis. The district should include comparable classrooms or schools, document how often teachers and pupils use the tool, and collect evidence from students, families, teachers, specialists, and administrators. Claims should be tested across age, disability, language background, and prior attainment so that average improvement does not conceal poor outcomes for a smaller group.
The strongest design compares AI-supported instruction with the normal instructional approach, not merely satisfaction before and after implementation. Even that design has limits because schools, teachers, and students may behave differently during a pilot. Therefore, the purchaser should ask whether the vendor supplied independent research, whether the study used the same product and version, and whether outcomes were reported by subgroup. “Personalized” does not necessarily mean effective; adaptation must change the learner’s task, feedback, timing, or support in a way that improves learning. Expansion should occur only when there is credible evidence of benefit and no serious unresolved safety concern.
What Should the AI Tutor Contract and Pricing Model Include?\n
Contracts should convert sales promises into enforceable service standards. The agreement should state the permitted educational uses, prohibited uses, student age range, supported languages, accessibility standard, data ownership, model-training restrictions, incident-response duties, audit rights, service levels, and termination rights. The school should know whether the tutor generates final answers, provides hints, corrects reasoning, or produces individualized lesson plans. It should also know which outputs require teacher review and which information is automatically transmitted to a third party. Schools should avoid vague terms such as “reasonable security” unless the document is supported by measurable standards and audit evidence.
Pricing should be compared on total cost rather than licence cost alone. Per-student annual pricing can appear economical, but districts may need teacher dashboards, administrator training, professional-development sessions, device purchases, speech services, accessibility support, migration, and additional storage. A budget based on all enrolled pupils may overstate actual use, while a model based only on active users can make a growing rollout harder to forecast. School buyers should request a first-year cost for 100, 500, and 1,000 students, plus quotes for training and support, so that scaling assumptions are visible. They should also determine whether a contract auto-renews, how notice periods work, and what happens to student records after termination.
The vendor should provide outcome-linked remedies, not only product credits. Appropriate remedies may include extension of the subscription, refund of unused fees, remediation, deletion of improperly processed data, or transition assistance. The procurement team should not accept a clause stating that the platform is educational or that the vendor is responsible for compliance if those promises impose no operational duty. A limited pilot agreement can reduce financial exposure, but it should use the same data, safety, accessibility, and deletion conditions expected in a full deployment. Savings claimed in later years should be considered hypothetical until the school has demonstrated sustained implementation and measured learning.
How Do AI Tutors Compare with Teachers, Human Tutors, and Other Digital Tools?
No AI tutor can be treated as a direct substitute for every part of teaching. Human tutors can diagnose motivation, interpret nonverbal cues, adapt to trauma or disability, model dialogue, notice emotional distress, and provide accountability in ways that a chatbot may not reproduce reliably. That does not make human tutoring a simple alternative for every school, because trained tutors may be unavailable, expensive, or difficult to recruit. AI can potentially provide inexpensive, consistently available practice, but only within a defined instructional design and with human access. Comparing options by asking what each one can do responsibly is more useful than declaring one category universally superior.
| Feature | AI tutor | Live human tutor | Conventional digital practice |
|---|---|---|---|
| Availability | Often available outside class and at consistent times | Depends on staffing and scheduling | Usually available whenever the platform works |
| Response consistency | Can provide repeatable feedback, but may generate errors | Can adjust live to learner behavior | Feedback varies by program and item type |
| Emotional support | Must be limited and escalated appropriately | Can offer nuanced social and emotional support | Usually limited to instructional prompts |
| Cost structure | Often per-student licences plus implementation costs | Usually highest hourly or programme cost | Often lower cost with limited personalization |
| Evidence requirement | Pilot results should be independently credible | Human-tutoring evidence exists, but fit depends on programme quality | Varies considerably by assessment design |
| Main procurement risk | Data, unsafe output, weak learning, opaque automation | Staffing, access, quality assurance, continuity | Low engagement, weak feedback, or teaching to the tool |
What Common Procurement Mistakes Should Districts Avoid?\n
One common mistake is selecting the vendor before defining the problem. Demonstrations are designed to show smooth conversation and instant feedback, which can obscure poor curricular alignment, hallucinated content, excessive answer-giving, or weak reporting. Another error is equating time on platform with learning; without a valid assessment, a highly engaged student may still have misunderstood the core concept. Districts should also avoid assuming that teachers will absorb training and monitoring work without scheduled time. If the tool requires teachers to review every interaction, that operational burden may exceed the advertised student licence saving.
Marketing claims require particular scrutiny because terms such as “adaptive,” “personalized,” and “evidence based” have no single universal meaning. Buyers should ask for the exact study, sample size, age range, duration, outcome measure, comparison group, and conflicts of interest. A five-question satisfaction survey or a small voluntary classroom trial is not enough to establish a causal effect. The California proposal to prohibit AI from acting as a teacher, and public debate in New York and elsewhere, show that the boundary between instructional support and unauthorized instructional authority is unsettled. Districts should monitor legal and policy developments through September 2026 and avoid contracts that depend on the tutor exercising powers the school is prohibited from delegating.
A further mistake is buying an enterprise platform for a narrow need. A district may pay more for an AI tutor than for a vetted maths practice service, curriculum-aligned homework support, or trained adult volunteers. Pilot use should be proportionate, reversible, and connected to a person responsible for following up. Finally, schools should not deploy a system before testing crisis language, offensive content, accessibility, multilingual performance, and behavior with real but controlled student scenarios. Privacy screening alone is not a safety evaluation.
When Should a District Act, Pilot, Pause, or Stop?
A district should act now to define policy and evaluate carefully, but it does not need to rush into a networkwide purchase. The immediate tasks are to inventory existing AI tools, identify which are already generating or influencing student content, establish an approval route, and ask legal and information-security teams to review the current contracts. Staff should be told that the appearance of an AI feature does not mean it has been approved. By the 2026-27 school year, a district should be able to name a responsible owner, approved data purposes, escalation procedures, and a record of authorized tools.
A pilot is appropriate when the use case is bounded, the risk is manageable, and there is a credible reason the product may help. The school should pause if the vendor cannot answer basic questions about data retention, third-party processing, model training, age restrictions, or incident response. It should stop a deployment if students receive dangerous advice, confidential information is exposed, material discrimination appears, teachers cannot review outputs, or data is collected beyond the stated purpose. A failed pilot should not automatically mean AI tutors never work; it may mean the product, use case, student group, implementation, or comparison method was unsuitable.
The decision to scale should be made at a scheduled review point rather than under sales or public-pressure pressure. Expansion should require acceptable learning results, stable costs, teacher usability, accessibility, complaint resolution, and documented safeguards. If those conditions are not met, the district should use the evidence to select a different tool or a non-AI alternative. This is not anti-technology. It is a recognition that responsible procurement protects instructional credibility and makes future innovation more sustainable.
What Questions Should Be Asked During Vendor Demonstrations?
Before a demonstration, the school should send the vendor a written scenario set based on the intended curriculum and student population. The vendor should then show, rather than merely describe, how the tutor handles an incorrect explanation, an uncertain fact, a request for personal data, a disability-related accommodation, a sensitive disclosure, and an exercise where the learner should think rather than receive the answer. A scripted conversation with a cooperative student is weak evidence. The demonstration should also show how teachers receive alerts, how administrators inspect appropriate activity records, and how data is exported or deleted.
The buyer should ask how often the underlying model changes and whether a previously validated version can remain stable for a school year. If the system is altered materially during the contract, the school should be notified and allowed to reassess its approval. References should include independent schools that have used the same version long enough to measure outcomes. Vendors may provide success stories, but customer transformation claims do not substitute for controlled evidence. The Manhattan Institute’s warning that one approach does not fit all K-12 settings should inform the question: what happened in classrooms unlike ours, including with younger children and students with additional needs?
Finally, the buyer should establish a written decision rule before seeing the vendor’s preferred result. For example, the school might require statistically or educationally meaningful improvement against a credible baseline, completion of at least one full term, no unresolved high-severity safety event, and acceptable performance for priority student groups. Thresholds should reflect the stakes and available assessment tools rather than a universal percentage invented by the purchaser. A product that scores well with advanced pupils but underperforms for pupils needing support should not pass simply because its overall average is high.
How Can a School Move from Evaluation to a Responsible Purchase?\n
A defensible process has five connected stages: need definition, due diligence, pilot, review, and contract. The need definition should specify the subject, learner group, instructional activity, desired outcome, and prohibited uses. Due diligence should examine research, privacy, security, accessibility, curriculum fit, implementation demands, and financial terms. The pilot should establish a baseline, train participating staff, monitor safety, and collect learning evidence from more than one viewpoint. Review should compare actual results with the criteria agreed before deployment. The final contract should preserve those protections after purchase.
The school should keep a record of why the product was selected, which alternatives were considered, who approved the pilot, what happened, and what remains uncertain. This makes the decision explainable to parents, teachers, governing bodies, auditors, and the next procurement team. It also reduces the risk of buying a new tool simply because the previous contract ended. Evidence from the UK National Tutoring Programme deployment can inform questions about tutoring design and use, but schools should not transfer results from one programme to another without examining the learner population, tutoring model, curriculum, and outcome measures.
The most authoritative answer is therefore procedural: procure an AI tutor only when the instructional need, student protections, human oversight, measurable learning benefit, and lifetime cost are all documented. As of 28 September 2026, schools should favor products that are transparent about limitations and teachers who retain authority over instruction. AI can provide useful additional practice and responsive feedback, but it should not receive authority merely because it is available around the clock. A school that follows that standard is not slowing innovation; it is making innovation accountable to the children it is intended to serve.