
How Long Does It Take to Build an EdTech PlatformRead More

Scaling an EdTech platform to millions of learners often requires help from an experienced in-house platform team, a specialist scalability consultancy, or an EdTech software development partner with expertise across architecture, infrastructure, data, AI, integrations, and DevOps.
The engineering work starts with real usage patterns. Millions of registered learners may translate into a much smaller peak of concurrent users, yet exams, enrollments, live sessions, or product launches can create sudden traffic spikes. Different EdTech products also stress different parts of the system, from database writes and search to video delivery, mobile sync, or AI inference.
This guide explains how to scale an EdTech platform through capacity planning, caching, database scaling, queues, auto-scaling, multi-region design, observability, load testing, disaster recovery, security, and cost control.
"EdTech platform" in this guide means the whole range of education software: LMS and LXP products, assessment and certification systems, AI tutors, tutoring and course marketplaces, mobile learning apps, digital curriculum and content platforms, authoring tools, learning analytics products, virtual classrooms, virtual labs, and workforce learning platforms. The core scaling principles remain consistent across all EdTech products, but each product creates different workload patterns, making different parts of the architecture the primary bottlenecks.
| Product type | Where scaling pressure usually lands |
|---|---|
| Assessment and certification | Writes, concurrency at start time, and zero data loss |
| AI tutors and assistants | Inference latency and cost per request |
| Content and curriculum platforms | CDN, media delivery, and search |
| Tutoring marketplaces | Matching, scheduling, and real-time communication |
| B2B learning SaaS | Multi-tenancy, role-based access, and integrations |
| Mobile learning apps | Offline sync and behavior on weak networks |
Registered learners is a marketing number. Concurrent learners is an engineering number, and infrastructure is sized from the second one.
Most registered users are inactive on any given day. Of those who are active, most arrive in a few busy hours, and each stays for a session of limited length. The number that loads the servers is how many sessions are open at once, multiplied by how many requests each session sends.
A platform with two million accounts and a platform with two hundred thousand accounts can need similar infrastructure if their busiest five minutes look alike. Platforms that size infrastructure for registered users overspend for months and still fall over on the one morning that matters, because the fixed capacity was never tested against the real peak.
Education traffic follows a calendar, which makes it unusually spiky and also unusually predictable. The common peaks are:
Each spike stresses a different part of the system. A term start is read-heavy and mostly cacheable. An exam is write-heavy and cannot lose a single answer, even when thousands of autosaves arrive in the same second. A live class barely touches the database but saturates media servers. Capacity planning works best when each of these scenarios is modeled separately.
A capacity model turns a business number (registered learners) into an engineering number (requests per second each component must handle). The chain has five steps:
Every input should come from the platform's own analytics where they exist. The worked example below uses assumed inputs to show the method.
| Step | Assumption | Result |
|---|---|---|
| Daily active learners | 10% of 2,000,000 registered | 200,000 |
| Busiest hour | 20% of the day's active learners | 40,000 |
| Peak concurrent users | 40,000 × 30-minute sessions ÷ 60 | 20,000 |
| Peak requests per second | 20,000 × 6 requests per minute ÷ 60 | 2,000 RPS |
| Design load | 2,000 × 2 headroom | 4,000 RPS |
| Application instances | 4,000 ÷ 200 RPS per instance at target latency | 20 instances |
| Database reads | 70% of requests are reads, 90% served from cache | 280 reads/sec reach the database |
| Database writes | 30% of requests are writes (progress, answers, events), batched by a queue in groups of 50 | about 24 batch writes/sec |
The two million registered learners become 20,000 concurrent users, 4,000 requests per second at design load, and a few hundred database operations per second once caching and queuing do their job. That is a very different system from one sized for two million simultaneous users.
The example also shows where the risk sits. If the cache hit rate falls from 90% to 50%, database reads jump from 280 to 1,400 per second. If writes go straight to the database without a queue, the database takes 1,200 writes per second. Connection limits matter too: 20 instances with a pool of 10 connections each means 200 database connections, so a connection pooler belongs in the design before autoscaling adds more instances.
Now take 100,000 of those learners sitting a timed exam at 9:00 AM. Assume 80% start within three minutes and each start sends about eight requests (sign-in, exam load, first questions): roughly 3,600 requests per second in the first minutes, from a much smaller audience than the daily peak. Autosaving every 30 seconds then produces about 3,300 writes per second for the length of the exam, and if most learners submit in the last two minutes, the system takes around 800 submissions per second at the end.
The daily model is read-heavy and cache-friendly. The exam model is write-heavy, and every write must survive. Platforms that only model the average day tend to discover the exam model in production.
Arbisoft's engagement with TenMarks, an Amazon company offering math practice used in more than 85% of U.S. school districts, covered web, iOS, and Android apps for more than 2 million users and over 220 million server requests per week, with server response times under 100 milliseconds. Spread evenly across all 168 hours of a week, 220 million requests is about 364 per second. A classroom product doesn't spread evenly. If most of that traffic fell inside roughly 35 school hours a week, the average during those hours would be closer to 1,700 requests per second, and the busiest minutes higher still. Weekly totals are useful for cost. Capacity has to be planned from the peak minute.
Each layer below exists to keep work away from the slowest and most expensive part of the system, which is usually the primary database.
Static assets, course images, SCORM packages, PDFs, and pre-rendered public pages should be served from a CDN close to the learner. A cached course catalog served from the edge never reaches the application. For platforms with learners across several countries, the CDN is also the cheapest way to cut latency for people far from the main region.
Video is usually the largest share of bytes a learning platform serves. The standard pattern is object storage for source files, transcoding into several renditions, adaptive bitrate streaming (HLS or DASH) so a learner on a weak connection gets a lower quality instead of a stall, and CDN delivery with signed URLs for paid content. Application servers should never stream video themselves. A managed video platform is usually the right call unless video is the product's differentiator.
Course structures, catalog listings, permission lookups, and configuration are read constantly and change rarely. An in-memory cache such as Redis in front of the database takes most of that load. The design question is invalidation: clearing a course's cache when it is published works better than a short expiry timer, which either serves stale content or misses too often.
Learning data is relational: learners, enrollments, attempts, grades, and certificates need transactions and constraints, so a relational database such as PostgreSQL remains a strong system of record at millions of users. Scaling it follows a predictable order. Fix query patterns and indexes first, because an N+1 query behaves badly at any size. Add read replicas for reporting and read-heavy pages, keeping in mind that replicas lag slightly, so a learner should read their own just-submitted answer from the primary. Partition large event tables by time. Move analytics to a separate store so an administrator's year-long report never competes with a learner's quiz submission. Splitting data across multiple primary databases (sharding, or one database per large tenant) comes last and only when a single primary can no longer take the write load.
Anything that doesn't need to finish before the learner sees a response belongs in a queue: certificate generation, grade passback to an LMS, notification emails, bulk enrollment imports, report exports, xAPI event forwarding, and AI-generated feedback. Workers process the queue at a controlled rate, so a spike becomes a backlog that drains in minutes instead of an outage. Every queued job needs to be safe to run twice, because retries happen.
Stateless application servers can scale horizontally behind a load balancer, which means session state, uploads, and caches cannot live on individual servers. Auto-scaling on CPU alone reacts too late for an exam that starts at a known minute, so scheduled scaling ahead of predictable peaks works better for education traffic. Auto-scaling also has a ceiling set by the database: thousands of new instances opening connections can exhaust it before the application tier notices, which is why connection pooling and queued writes come first.
Course marketplaces, content libraries, and curriculum products with large catalogs eventually need search as its own system. A separate search index (Elasticsearch or OpenSearch) handles faceting, relevance tuning, and typeahead without loading the primary database. An indexing pipeline keeps it updated from change events, which means search results are eventually consistent, usually seconds behind a publish. Recommendation workloads belong alongside search, computed in the background and served from cache.
Mobile learning apps reach many of the learners that drive growth in global education, often on mid-range devices and intermittent connections. Scaling them is mostly about the API and sync design. Offline learning means storing lessons and progress on the device and syncing when a connection returns, with conflict rules decided up front. Every write from the app should carry an idempotency key so retries on a flaky network never create duplicate attempts or submissions. Compact API responses, batched requests, and low-bitrate media renditions cut both data use and server load. Old app versions stay in use for months, so APIs need versioning and backward compatibility.
Large EdTech ecosystems connect to SIS and LMS platforms, HRIS and CRM systems, identity providers, payment gateways, proctoring services, content providers, LTI tools, xAPI learning record stores, and external AI services. At millions of learners, those connections become bottlenecks of their own: providers enforce rate limits, slow down during their own peaks, and send webhooks in bursts. Calls to them should go through queues, with retries and backoff, circuit breakers that stop calling a failing service, and webhook handlers that accept the event quickly and process it later. Each integration needs its own monitoring, because a learner sees a failed grade sync as a platform bug.
Many EdTech platforms reach millions of learners as hundreds of institutions or enterprise customers with tens of thousands of learners each. That changes the scaling problem. One large tenant running a year-end report or a district-wide exam can slow everyone else, so tenant-aware rate limits, per-tenant queues for heavy jobs, and per-tenant metrics belong in the design. Most platforms start with shared tables and a tenant identifier enforced in every query, then move their largest or most regulated customers to dedicated databases when isolation or residency requires it. Deciding this early matters because the tenancy model shapes the data model, and changing it once customer data exists is slow and risky.
A second region earns its cost when learners in another geography need lower latency, when data residency rules require local storage, or when the business cannot accept a regional cloud outage. Scale across regions also means more devices, network conditions, and learner needs, so accessibility and performance on low-end devices should be tested in every market the platform enters. The simplest working pattern is one primary region for writes, read replicas and CDN caching closer to learners, and a warm standby that can take over. Active-active writes across regions add conflict handling that most learning platforms don't need.
At millions of learners, failures are certain and the goal is to find them fast, contain them, and recover in a planned way.
Logs, metrics, and traces should answer three questions within minutes: what is slow, for whom, and since when. The metrics that matter most in EdTech are tied to learner actions: login success rate, time to load a lesson, answer-save latency, submission success rate, and queue depth. Alerts belong on those, with thresholds set from real peaks.
Load tests should replay the scenarios from the capacity model, including the exam morning, at the design load and beyond it, to find the breaking point before learners do. A university that will run an exam for 10,000 students should test 10,000 simultaneous starts, with realistic think time and autosave behavior. Testing a single endpoint in isolation misses the contention that appears when everything happens at once.
Some features can fail without harming learning, and the system should know which ones in advance. Recommendations, leaderboards, activity feeds, and analytics widgets can switch off under pressure. Content delivery, answer saving, and submission cannot. Feature flags and per-feature timeouts make that decision automatic, so a slow recommendation service returns nothing instead of blocking a lesson page.
More learners means more attack surface and more automated abuse. Rate limiting on sign-in, password reset, and public APIs stops credential stuffing and bots before they reach the application. A web application firewall and DDoS protection at the CDN edge absorb traffic the origin should never see. Authentication itself can become a bottleneck on exam mornings, so identity services need the same load testing as the rest of the platform. Tenant isolation, secrets kept in a managed vault, and audit logs for grade changes and admin actions are what enterprise and institutional buyers will ask to see.
Backups only count once a restore has been tested. Each system needs a recovery point objective (how much data the business can lose) and a recovery time objective (how long it can be down), and those targets are much stricter for an exam platform than for a course catalog. Restores and regional failover should be rehearsed on a schedule, not attempted for the first time during an incident.
Infrastructure cost should grow more slowly than learner numbers. The useful metric is cost per monthly active learner, tracked by component: compute, database, storage, CDN and video, and third-party services.
The biggest savings usually come from the same places as the performance gains. A higher cache hit rate means smaller databases. Queues let workers run on cheaper capacity. Scheduled scaling avoids paying for peak capacity all week. Video transcoding settings and storage tiers can cut the largest line item. Committed-use discounts make sense for the steady baseline, with on-demand capacity covering spikes.
AI tutors, feedback, and authoring tools scale differently from page traffic. Each request can take seconds and costs money per token, so one million AI-assisted learners is a cost and latency problem before it is a server problem. Platforms that handle this well route requests through a gateway that tracks cost per tenant, cache answers to repeated questions, use smaller models where quality allows, generate content asynchronously when no learner is waiting, and set a timeout with a fallback for every model call. Retrieval for grounded answers needs its own capacity planning, since vector search load grows with both content volume and query traffic.
The right help comes from an engineering team that has already operated an education platform at the scale you are planning for and can show the evidence. Four kinds of teams usually do this work:
The best fit depends on what is breaking. A single slow database is an optimization project. A platform that needs multi-tenancy, a second region, an exam engine, and lower cost per learner at the same time is an architecture and product project, and it needs a team that understands education workflows as well as infrastructure.
Whoever is shortlisted, a few questions separate real experience from a sales deck. Ask for the peak concurrency of a platform they have run, and what broke first. Ask how they would load test your exam morning. Ask to see an observability dashboard or incident review from a past project. Ask how they would lower cost per active learner. The guide on how to choose an EdTech development company has the wider evaluation checklist.
Arbisoft has been part of this work since 2013. The edX engagement grew to a team of more than 150 experts working on a platform used by more than 20 million learners and 140 partners, covering course tooling, proctored exams, analytics, enterprise SSO, and LMS integrations. Arbisoft's EdTech division also maintains Tutor, the Docker-based distribution the Open edX documentation recommends for installing and running the platform. The TenMarks work described above added consumer-scale classroom traffic across web and mobile. For the stack choices behind each layer, the guide to the best tech stack for a scalable learning platform goes deeper.
It depends on daily activity, peak-hour concentration, and session length. Using the model above, 1 million registered learners with 10% daily activity, 20% of that in the busiest hour, and 30-minute sessions gives about 10,000 concurrent users. Exams and deadlines can multiply that for short windows, so model those separately.
Usually later than teams expect. A well-structured modular monolith scales to millions of learners when caching, queues, and the database are designed well. Separate services earn their cost when a component needs to scale on its own, such as video processing, exam delivery, search, or AI inference, or when teams block each other on one deployment.
Open edX can, and it runs edX, which serves more than 20 million learners. Getting there takes the same work as any platform: caching, separate workers for background tasks, database tuning, CDN delivery, and load testing against real peaks. Customizations need to stay close to the upstream code so upgrades remain manageable.
The database is the most common first failure, often through connection exhaustion or slow queries that were harmless at small scale. Reports run against production tables, video served from application servers, and synchronous calls to third-party services are the next most common causes.
Build the capacity model for your two hardest scenarios, the busiest normal hour and the biggest exam or launch, using your own analytics. Then load test against both. The gap between the result and the design load tells you exactly which layer to fix first, and it gives any partner you bring in a concrete brief to work from.
Trusted by top platforms for our transformative solutions and exceptional results:






