Scalability Model
Scalability Overview
PeopleHub is designed to scale from 100 employees to 10,000+ without architectural changes, leveraging AWS serverless auto-scaling capabilities.
Component Scalability
Frontend (S3 + CloudFront)
Current Capacity: Unlimited Scaling Mechanism: Automatic (AWS-managed)
CloudFront edge locations automatically serve static assets globally:
- 200+ edge locations worldwide
- Automatic cache scaling
- No configuration needed
Performance at Scale:
- 100 users: <50ms latency
- 10,000 users: <50ms latency (same)
- No difference in frontend performance regardless of user count
API Gateway
Current Capacity: 10,000 requests/second (regional default) Scaling Mechanism: Automatic
- Burst capacity: 5,000 requests
- Sustained capacity: 10,000 requests/second
- Can be increased via AWS support request (no limit)
No action required for scaling up to 10,000 concurrent users.
Lambda Functions
Current Capacity: 1,000 concurrent executions (default) Scaling Mechanism: Automatic, sub-second
How Lambda Scales:
- First request creates new Lambda instance
- Request processed (warm execution: 10-50ms)
- Subsequent requests reuse warm instance
- If all instances busy, AWS creates new instances automatically
- Scales from 0 to 1,000 concurrent executions in seconds
Scaling Example:
- Normal: 50 concurrent requests → 50 Lambda instances
- Peak: 500 concurrent requests → 500 Lambda instances (auto-scaled)
- Night: 0 requests → 0 Lambda instances (scaled to zero, no cost)
Limits:
- Default: 1,000 concurrent executions per region
- Can be increased to 10,000+ via AWS support
- No application changes needed
RDS PostgreSQL
Current Configuration: Single instance, Multi-AZ Scaling Mechanism: Vertical scaling (manual), read replicas (future)
Vertical Scaling:
- Current: t3.medium (2 vCPU, 4 GB RAM) - sufficient for 1,000 employees
- Scale up: t3.large, t3.xlarge, or m5/r5 instances
- Scaling requires brief downtime (few minutes) for instance change
Connection Pool:
- Each Lambda maintains 10 connections
- RDS supports 100-1000+ connections based on instance size
- RDS Proxy can be added for better connection management (if needed)
Read Replicas (future):
- Create read-only replicas for reporting
- Offload read traffic from primary database
- Up to 15 read replicas supported
S3 Document Storage
Current Capacity: Unlimited Scaling Mechanism: Automatic
S3 automatically scales to handle:
- Unlimited files
- Unlimited storage
- Thousands of requests per second
No configuration needed for scaling.
Scaling Scenarios
Scenario 1: Company Growth (100 → 1,000 Employees)
Frontend: No changes needed Lambda: Auto-scales to handle increased API traffic RDS: May need to scale up instance size S3: No changes needed
Action Required: Monitor RDS performance, scale up if needed
Scenario 2: High-Volume Onboarding (50+ Candidates/Month)
Candidate API Lambda: Auto-scales to handle concurrent onboarding S3: Auto-scales for document uploads Notifications: Auto-scales for email sends
Action Required: None, automatic scaling handles this
Scenario 3: Global Expansion (7 Offices Worldwide)
CloudFront: Already global, no changes RDS: Add read replicas in other regions Lambda: Deploy to multiple regions (optional for reduced latency)
Action Required:
- Setup cross-region RDS read replicas
- Configure Route 53 geolocation routing (optional)
Scenario 4: Peak Usage (All Employees Access Simultaneously)
Example: Company announcement at 9 AM, 5,000 employees log in
Traditional System:
- Servers overload
- 503 errors
- 5-10 minutes to auto-scale
PeopleHub Serverless:
- Lambda scales to 5,000 concurrent executions in <2 seconds
- CloudFront serves frontend from cache (no backend load)
- All users served successfully
- No manual intervention
Performance Targets by Scale
| Users | Concurrent Requests | Lambda Instances | API Latency | RDS Instance |
|---|---|---|---|---|
| 100 | 10 | 10 | <100ms | t3.small |
| 500 | 50 | 50 | <150ms | t3.medium |
| 1,000 | 100 | 100 | <200ms | t3.large |
| 5,000 | 500 | 500 | <200ms | m5.xlarge |
| 10,000 | 1,000 | 1,000 | <250ms | m5.2xlarge + read replicas |
Cost Scaling
Cost Comparison at Different Scales
At 100 Employees (Low Traffic):
- Traditional servers: $250/month (always running)
- PeopleHub serverless: $60/month (pay-per-use)
- Savings: 76%
At 1,000 Employees (Moderate Traffic):
- Traditional servers: $600/month
- PeopleHub serverless: $180/month
- Savings: 70%
At 5,000 Employees (High Traffic):
- Traditional servers: $2,000/month
- PeopleHub serverless: $600/month
- Savings: 70%
Key Insight: Serverless becomes MORE cost-effective at scale due to efficient resource utilization.
Scaling Limits & Mitigation
Lambda Concurrency Limit
Default: 1,000 concurrent executions Mitigation: Request increase via AWS support (can scale to 10,000+) When Needed: At 5,000+ concurrent users
RDS Connection Pool
Limit: Based on instance size (e.g., 500 connections for m5.large) Mitigation: Add RDS Proxy for connection pooling When Needed: At 2,000+ concurrent database operations
API Gateway Throttling
Default: 10,000 requests/second Mitigation: Request quota increase via AWS support When Needed: At 10,000+ sustained requests/second
Monitoring for Scalability
Key Metrics to Watch:
- Lambda concurrent executions (alert at 80% of limit)
- RDS CPU and connections (alert at 70% utilization)
- API Gateway 4xx/5xx errors (alert on spike)
- Application response times (p95 latency)
CloudWatch Alarms (planned):
- Lambda throttling errors
- RDS connection exhaustion
- API Gateway throttling
- High error rates
Auto-Scaling vs Manual Scaling
| Component | Scaling Type | Configuration Needed |
|---|---|---|
| CloudFront | Auto | None |
| API Gateway | Auto | None |
| Lambda | Auto | None (monitor limits) |
| RDS | Manual | Instance size change |
| S3 | Auto | None |
Only RDS requires manual intervention for vertical scaling.
Future Scalability Enhancements
Multi-Region Deployment
- Deploy Lambda functions in multiple AWS regions
- Read replicas in each region
- Route 53 for geolocation-based routing
- Reduces latency for global users
Aurora Serverless v3
- Reevaluate Aurora Serverless (better than v2)
- Automatic database scaling (ACUs)
- No manual instance sizing
ElastiCache (Redis)
- Add caching layer for frequently accessed data
- Reduce database load
- Improve response times
SQS for Async Processing
- Queue long-running operations
- Process asynchronously via separate Lambda
- Improve API response times
Scalability Best Practices
- Stateless Design: No session state in Lambda (enables horizontal scaling)
- Database Connection Pooling: Reuse connections across requests
- Efficient Queries: Optimize SQL, use indexes, paginate results
- CDN Caching: Maximize CloudFront cache hit ratio
- Async Operations: Use queues for non-critical operations
- Monitoring: Proactive monitoring and alerting for limits
Related Documentation
- System Overview - Architecture diagram
- AWS Architecture - AWS services
- Performance Benchmarks - Performance data
- Monitoring - Observability