
How to Build an Adaptive Learning Platform with Real Time PersonalizationRead More

AI course authoring can accelerate course creation, but effective implementation requires more than adding an LLM to a learning platform.
A stronger approach is to build authoring around structured course data linked to trusted source documents. Content is generated in stages, with SMEs reviewing key decisions such as the course outline, while assessments, learner variants, translations, and updates all draw from the same underlying structure.
This guide explains how to build AI course authoring for a learning platform, from the underlying architecture and generation workflow to assessment design, quality controls, tooling choices, and implementation order.
Automated course content creation shortens the assembly work in production and leaves review time roughly where it was. In Chapman's breakdown for basic e-learning, authoring and programming took 17% of development hours and instructional design 14%, with storyboarding and graphics close behind. SME and stakeholder review took 7%, and quality assurance 6.5%.
Controlled trials of generative AI in edtech and adjacent writing work point the same way. A preregistered experiment published in Science.org with 453 professionals found ChatGPT cut time on writing tasks by 40% while rated quality rose 18%. An Education Endowment Foundation trial with 259 teachers in England found lesson preparation time fell 31%, from 81.5 to 56.2 minutes a week, and a blind expert panel found no drop in quality.
Review stays because models fail unevenly. In a field experiment with 758 BCG consultants, people using GPT-4 on tasks outside the model's competence were 19 percentage points less likely to reach a correct answer than those working without it. An authoring system has to assume some drafts fall outside that frontier and send them to a person who can tell.
A course graph is the data model that makes AI course authoring maintainable. Every learning objective, outline node, content block and assessment item is a versioned record, and each one links to the source passages it was generated from. The language model writes into the graph and never writes a finished package directly.
The reason is that every feature buyers ask for reads from the same records:
| Feature | What it reads from the course graph |
|---|---|
| AI-generated assessments | Objectives, their target cognitive level, and the blocks that teach them |
| Learner-segment variants | Objectives plus a segment profile (role, industry, reading level) |
| Multilingual course localization | Text fields of blocks and items, plus a locked termbase |
| Keeping content current | Provenance links from blocks and items back to source passages |
| Item analytics | Stable item IDs that survive edits and exports |
Without the graph, each of those becomes a separate project that re-parses exported HTML.
Structured outputs make the graph practical to fill. OpenAI and Anthropic both document schema-constrained generation that holds a response to a supplied JSON Schema. That constraint covers shape. A block can validate perfectly and still misstate a policy, which the next two components handle.
Many platforms already hold part of this model in their import and export code, where the implicit content model is written down. Arbisoft's work on the edX platform since 2013 included refining Studio's course import/export and content creation features, and that layer is the natural starting point for a graph.
Grounding means the generator writes each block from retrieved passages of the customer's own documents and records which passages it used. Ingestion splits source files (policies, product documentation, SME slide decks, call transcripts) into passages with stable IDs and keeps each file's version so later changes can be traced.
Retrieval improves factuality without guaranteeing it. In the original retrieval-augmented generation paper, human raters judged the retrieval model more factual than a plain generator in 42.7% of comparisons and less factual in 7.1%. A Stanford study of commercial legal research tools built on retrieval found they still hallucinated between 17% and 33% of the time.
So every block cites the passage IDs it drew on, and an automated check confirms each passage supports its claim before an SME sees the block. The retrieval layer is standard RAG pipeline engineering. The course-specific part is that citations point at passage records in the graph, so they survive editing and translation.
A staged pipeline generates a course in the order an instructional designer would plan it, and stops for human approval where one decision shapes the most content:
Steps 2 and 3 amount to an AI curriculum builder. The gate sits at step 4 because a scope change there costs a few minutes of editing, while the same change after step 6 means regenerating and re-reviewing dozens of blocks and items. An outline also asks the question non-expert SMEs are best placed to answer from experience: is this what people need to learn?
The review screen decides whether SMEs trust the tool. Each block appears beside the passages it cites, and every accept, edit or rejection is logged. That log becomes the main quality metric and, as the FAQ explains, part of the customer's copyright position.
Each stage should also be callable as an API, so the studio, LMS plug-ins and AI agents share one pipeline. Arbisoft's Edtech division, Edly, builds Compose, which lets creators author with AI directly inside LMSs including Canvas and Blackboard.
AI-generated assessments hold up when the generator works from an item blueprint and the platform treats every item as provisional until learner data confirms it. Three recent studies explain both conditions.
In a 2025 study in BMC Medical Education, GPT-4o wrote in 24.5 person-hours a question set that took humans 96. Its items were easier, and 84% sat at Bloom's Remember or Understand levels, while 44% of the human items reached Apply or Analyse. A blinded, preregistered study in npj Digital Medicine found that expert-reviewed GPT-4o items matched human items on difficulty and discrimination. Examinees couldn't tell them apart. A 2026 network meta-analysis of 15 studies found no significant quality gap for GPT-4 items and rated the certainty of all its evidence "very low".
Four design choices follow from that evidence. The blueprint assigns each objective a target cognitive level and an item count, and the generator is prompted per blueprint cell, which counters the drift toward recall questions. A similarity check removes near-duplicate items before review. New items stay unscored or low-weight until enough learner responses exist to estimate difficulty and discrimination, and items that discriminate poorly are flagged for retirement. Items export as QTI 3.0, the 1EdTech standard for moving tests and questions between systems.
AI-generated questions can be used in certification exams under process rules that accreditation and testing bodies have now written down. No peer-reviewed study yet reports how LLM-generated items performed on a live certification exam, so the rules define a defensible process without claiming equivalence.
The NCCA's guidance on AI in certification programs (April 2025) states that "SMEs must review all AI-generated content to ensure accuracy, fairness, and compliance with psychometric standards." It also requires programs to document AI's role and reserves final item decisions for qualified people. The ITC and ATP Guidelines for Technology-Based Assessment (version 1.1, July 2025) add fairness and bias reviews and state that "field testing is needed to gather data needed to support calibration of AIG items."
The Duolingo English Test is the clearest published example at high stakes. In its reading task pilot, 454 of 789 generated passages (58%) survived human review, and each passage received at least three content reviews and two fairness reviews. The listening task kept 713 of 900 (79%). The test's technical manual describes the operational pipeline: generation, human review, calibration, and regular checks for differential item functioning.
For a platform, this adds up to a certification mode: mandatory SME sign-off on every item, an audit log of what AI drafted and what people changed, pretest slots for new items, and exportable documentation for the customer's accreditor. EU customers need one more check. The AI Act lists AI systems "intended to be used to evaluate learning outcomes" as high-risk under Annex III, with those obligations applying from 2 December 2027 under the 2026 amendment. Whether an item-generation tool falls inside that definition is a question for counsel.
Personalized variants work when they are child records of the same objective and share its assessment items. A sales-engineer variant and a support-agent variant of a compliance lesson can use different examples and reading levels while learners answer the same calibrated questions, so scores stay comparable across segments.
Each variant adds review load, which argues for a small set of segment profiles per tenant. Choosing which variant a learner sees, and what comes next, belongs to the adaptive layer. The authoring system's job is to produce variants tagged with objectives and item IDs that the adaptive layer can route between, covered in our guide to adaptive learning platform development.
Multilingual course localization works best on the course graph itself. Translating block and item fields in the graph keeps provenance, item IDs and calibration data attached to every language version, and a locked termbase keeps product names and regulated terms consistent. A translation memory avoids paying to retranslate unchanged text after each update.
Model choice needs testing per language pair. In the WMT25 shared task, professional annotators scored English to Egyptian Arabic translations: human translators reached 78.5 out of 100 and the best system, GPT-4.1, reached 77.0, while other major LLMs scored between 55.7 and 74.0. The organizers found systems "tend to output Modern Standard Arabic" when Egyptian Arabic was requested, an error quality estimation metrics would likely miss, and concluded that "human evaluation should remain the final arbiter." Whichever variety a customer wants, the platform should request it explicitly and have people check it. For published courses, the safe default is what ISO 18587 calls full post-editing: human review until the output is "comparable to a product obtained by human translation."
Arabic also needs right-to-left rendering. W3C guidance is to set dir="rtl" on the html element, use logical CSS values (start and end in place of left and right), and wrap inserted text such as learner names in bdi elements. Digits follow the market, and the W3C's Arabic layout requirements list Saudi Arabia and Egypt among the countries using Arabic-Indic digits.
In Quebec, French is a legal requirement. The Charter of the French Language, as amended by Bill 96, requires employers to provide "training documents produced for the staff" in French on terms at least as favourable as any other language, and requires computer software to be available in French unless no French version exists. A B2B platform selling into Quebec needs French authoring screens and French course output.
Automated gates catch the failures that are cheap to detect before a person spends time on them: schema validity, citation support, termbase compliance, reading level, duplicate items, and accessibility problems such as missing alt text. The expensive check is semantic (does a block say what its sources say?), and the practical option at volume is a second model acting as judge.
An LLM judge is usable with safeguards. The MT-Bench study found GPT-4's judgments agreed with human experts over 80% of the time, the same rate at which humans agreed with each other. The same study found GPT-4 kept its verdict only 65% of the time when two answers were swapped in order. On math questions it failed to catch wrong answers 14 times out of 20 with a default prompt, and 3 times out of 20 when given a reference answer. For course content the reference is the cited source passage, so the judge always receives it. Answer order is randomized, and the judge is calibrated against a set of SME-labeled blocks before its scores gate anything.
Provenance links turn content freshness into a targeted job. When a customer uploads a revised policy, the platform compares it with the indexed version, finds the changed passages, and marks every block and item that cites them as stale. Only those are regenerated and re-reviewed. The rest keeps its approvals and item statistics.
| Component | Decision | What to look for |
|---|---|---|
| Model access | A gateway between the pipeline and model providers | Per-tenant routing rules, cost tracking, fallback models |
| Generation format | Schema-constrained structured outputs | Support for the schema features the graph uses |
| Retrieval | Vector plus keyword search over passage records | Tenant isolation, passage IDs, versioned indexes |
| Orchestration | A durable workflow engine | Long-running jobs, retries, human review as a workflow state |
| Evaluation | Judge model plus an SME-labeled test set | Versioned prompts, scores tracked per release |
| Packaging | SCORM, cmi5, LTI 1.3, QTI 3.0 | Every export generated from the graph |
ADL now lists SCORM, which it created in 2000, among its past projects. Its alternative, cmi5, is a profile of xAPI (standardized as IEEE 9274.1.1-2023) that adds rules for launch, sessions and course structure. LTI 1.3 with Deep Linking lets an instructor pull generated content into a course from inside the customer's own LMS.
Data residency narrows model choice. As of September 2026, the providers' availability tables show few frontier models running inside GCC regions:
The rules differ by market. Saudi Arabia's Personal Data Protection Law sets conditions on transfers abroad, including adequate protection and limiting data to the minimum needed, and SDAIA's transfer regulation requires a risk assessment in some cases. Quebec's Law 25 requires a privacy impact assessment before personal information leaves the province. Federal PIPEDA guidance permits transfers but relies on contracts for protection.
Course source documents often contain personal data, from named case studies to support transcripts. The model gateway therefore needs a residency policy per tenant, with an open-weight model hosted in-region as the fallback for tenants whose data can't leave. Provider tables change often, so recheck them before committing.
| Release | Ships | Proves |
|---|---|---|
| 1 | Source ingestion, course graph, staged pipeline with outline gate, practice quizzes, SCORM and cmi5 export, edit logging | SMEs publish a usable course faster than before |
| 2 | Citation checks, calibrated judge, item calibration from learner data, QTI export, source-change detection, Arabic and Canadian French | Quality holds as volume grows |
| 3 | Segment variants, adaptive hand-off, certification mode, per-tenant residency routing | The feature wins deals with regulated and GCC buyers |
The build-or-integrate decision comes down to the graph. An external authoring tool connected by LTI or SCORM reaches customers sooner, but its content lives outside the platform's data model, so freshness tracking, variants and item analytics can't reach it. Integration suits platforms treating authoring as a checklist item, and building suits those planning to compete on it.
Use retrieval for facts and keep fine-tuning for style. Course facts come from each customer's documents and change often, so they belong in an index that can be updated and cited. Fine-tuning can teach a house tone or item format, but it can't cite sources, and each base-model upgrade means retraining. In our view, retrieval plus structured prompts covers what most platforms need.
In the US, human contributions to AI-assisted content can be protected and purely AI-generated material cannot. The US Copyright Office concluded in January 2025 that "prompts alone do not provide sufficient human control" to make someone the author, while creative selection, arrangement and modification of AI output can be protected. SME edits and outline decisions matter, and the edit log is evidence of them. Customers should take legal advice for their jurisdiction.
Track time from source upload to published course, the share of generated blocks SMEs accept without edits, pass rates at each review gate, and item statistics once learners respond. Edit rate by block type shows where the pipeline is weak: if scenario blocks get rewritten twice as often as definitions, the scenario prompt needs work. The baseline should be the platform's own pre-AI numbers, since the most cited benchmark dates from 2010.
Build the course graph and the outline gate before comparing models, because both survive every model change. Then run one real course through the pipeline with one customer's SMEs and measure the edit rate, which tells the team which component to improve next. Arbisoft's education team has worked on learning platforms with edX since 2013 and builds AI course authoring through its Edly division.
Trusted by top platforms for our transformative solutions and exceptional results:






