# Scheduler / Queue / Runtime — Droplet Operations Reference

Companion to `doc/background works/existing issue with fixing.md`. Covers the server-side
configuration that the application code cannot set itself. The code changes that go with this doc
(schedule consolidation, `scheduled_task_runs` monitoring, `hr:mark-absent-staff` and
`jwt:generate-secrets` hardening) are already in the repo — see the **Deploy** section.

All deployments run the **same** `IMS-Backend-New` codebase, one directory + database + `.env` per
client:

```
api_gsc_backend   api_ims_demo_backend   api_mcas_backend   api_nextgen
api_pche_backend  api_pebbles_backend    api_tbs           api_theskillscp_backend
```

Runtime facts (from `.env`): `QUEUE_CONNECTION=database`, `CACHE_STORE=file`,
`SESSION_DRIVER=database`, `DB_CONNECTION=pgsql`, `APP_TIMEZONE=Asia/Colombo`. Because the queue,
cache, and sessions are all per-database / per-`storage`, **clients are already isolated** — there is
no shared Redis namespace or shared queue to untangle. The one hard rule: each deployment must keep
its **own `storage/` directory** (the file-cache-backed scheduler mutex and the API response cache
live there).

---

## 1. Schedule inventory — before vs after

### Before (two registration points, conflicting frequencies)

| Command | `routes/console.php` | `bootstrap/app.php` `withSchedule()` | Effect |
|---|---|---|---|
| `attendance:auto-checkout` | every minute | daily 17:00 | ran **both** |
| `commissions:close-expired` | hourly | daily 00:00 | ran **both** |
| `facebook:record-webhook-health` | every 15 min | hourly | ran **both** |
| `students:check-payment-status` | daily 00:00 | daily 00:00 | ran **twice at 00:00** |
| `hr:mark-absent-staff` | every minute | — | ok |
| `communications:process-scheduled` | every minute | — | ok |
| `app:prune-expired-file-cache` | hourly | — | ok |
| `hr:check-contracts` | daily 01:00 | — | ok |
| `backup:database-auto` | twice daily 00:00 / 12:00 | — | ok |
| `jwt:generate-secrets` | — | weekly Sun 02:00 | only in bootstrap |
| `facebook:process-webhook-retries` | — | every 5 min | only in bootstrap |
| `tiktok:sync-leads` | — | every 15 min | only in bootstrap |
| `tiktok:record-webhook-health` | — | hourly | only in bootstrap |
| `tiktok:process-webhook-retries` | — | every 5 min | only in bootstrap |
| `students:send-payment-reminders` | — | daily 08:00 | only in bootstrap |

### After (single canonical source: `routes/console.php`)

`bootstrap/app.php` `->withSchedule()` is removed. Every entry is registered once via
`App\Support\ScheduleMonitor::track()`, which pins `Asia/Colombo` and logs each run to
`scheduled_task_runs`.

| Command | Frequency | Cron | Why this frequency |
|---|---|---|---|
| `attendance:auto-checkout` | every minute | `* * * * *` | evaluates per-staff office-end / overtime cutoffs from `Setting`; needs minute granularity |
| `hr:mark-absent-staff` | every minute | `* * * * *` | self-gates on the per-client absent-threshold time |
| `communications:process-scheduled` | every minute | `* * * * *` | time-sensitive email/SMS fallback dispatcher |
| `facebook:process-webhook-retries` | every 5 min | `*/5 * * * *` | dispatches retry jobs; carried over from bootstrap |
| `tiktok:process-webhook-retries` | every 5 min | `*/5 * * * *` | carried over from bootstrap |
| `facebook:record-webhook-health` | every 15 min | `*/15 * * * *` | metrics snapshot; finer than the old hourly copy, kept |
| `tiktok:sync-leads` | every 15 min | `*/15 * * * *` | polling half of the hybrid lead ingestion |
| `app:prune-expired-file-cache` | hourly | `0 * * * *` | file-cache sweep; short, runs synchronously |
| `commissions:close-expired` | hourly | `0 * * * *` | hourly supersedes the old daily-midnight copy |
| `tiktok:record-webhook-health` | hourly | `0 * * * *` | carried over from bootstrap |
| `students:check-payment-status` | daily 00:00 | `0 0 * * *` | de-duplicated; now chunks + fans out to `EvaluateStudentPaymentStatusJob` (§4a) |
| `hr:check-contracts` | daily 01:00 | `0 1 * * *` | unchanged |
| `students:send-payment-reminders` | daily 08:00 | `0 8 * * *` | migrated; now chunks + fans out to `SendPaymentReminderBatchJob` (§4a) |
| `backup:database-auto` | 00:00 & 12:00 | `0 0,12 * * *` | unchanged |
| `jwt:generate-secrets` | weekly Sun 02:00 | `0 2 * * 0` | carried over; now also reloads runtime (see §6) |

**Nothing was removed or slowed down.** Four duplicates were resolved to the finer of the two
frequencies; six bootstrap-only entries were migrated verbatim.

### Verify on each deployment

```bash
for d in /var/www/html/api_*; do
  [ -f "$d/artisan" ] || continue
  echo "========== $d =========="
  (cd "$d" && php artisan schedule:list)
done
```

Every command above must appear **exactly once**. Then confirm the second registration point is gone:

```bash
grep -n "command(" /var/www/html/api_*/bootstrap/app.php   # expect: no matches
```

---

## 2. Cron — one trigger per deployment, with visible output

Replace the `>> /dev/null` entries. `www-data` crontab (`crontab -u www-data -e`):

```cron
* * * * * cd /var/www/html/api_gsc_backend        && php artisan schedule:run >> /var/log/laravel/api_gsc_schedule.log 2>&1
* * * * * cd /var/www/html/api_ims_demo_backend   && php artisan schedule:run >> /var/log/laravel/api_ims_demo_schedule.log 2>&1
* * * * * cd /var/www/html/api_mcas_backend       && php artisan schedule:run >> /var/log/laravel/api_mcas_schedule.log 2>&1
* * * * * cd /var/www/html/api_nextgen            && php artisan schedule:run >> /var/log/laravel/api_nextgen_schedule.log 2>&1
* * * * * cd /var/www/html/api_pche_backend       && php artisan schedule:run >> /var/log/laravel/api_pche_schedule.log 2>&1
* * * * * cd /var/www/html/api_pebbles_backend    && php artisan schedule:run >> /var/log/laravel/api_pebbles_schedule.log 2>&1
* * * * * cd /var/www/html/api_tbs                && php artisan schedule:run >> /var/log/laravel/api_tbs_schedule.log 2>&1
* * * * * cd /var/www/html/api_theskillscp_backend && php artisan schedule:run >> /var/log/laravel/api_theskillscp_schedule.log 2>&1
```

```bash
sudo mkdir -p /var/log/laravel && sudo chown www-data:www-data /var/log/laravel
```

Rotate them — `/etc/logrotate.d/laravel-schedule`:

```
/var/log/laravel/*_schedule.log {
    daily
    rotate 14
    missingok
    notifempty
    compress
    delaycompress
    copytruncate
    su www-data www-data
}
```

Also rotate the per-command app logs the scheduler writes (Laravel's own `storage/logs/laravel.log`
is already handled by `LOG_DAILY`, but add the schedule dir if you enable `appendOutputTo`):

```
/var/www/html/api_*/storage/logs/*.log {
    daily
    rotate 14
    missingok
    notifempty
    compress
    delaycompress
    copytruncate
}
```

### Confirm there is no second scheduler

```bash
grep -R "schedule:run" /etc/cron* /var/spool/cron* /var/www/html 2>/dev/null
systemctl list-timers --all | grep -i laravel
ps aux | grep '[s]chedule:run'
```

Expect only the eight cron lines above. No systemd timer, no duplicate crontab (root vs `www-data`),
no `schedule:run` inside any deploy script.

---

## 3. Monitoring — `scheduled_task_runs`

The consolidation adds a table (migration `2026_08_29_120000_create_scheduled_task_runs_table`) and
`App\Models\ScheduledTaskRun`. `ScheduleMonitor::track()` writes one row per completed run:
`command`, `started_at`, `finished_at`, `duration_ms`, `status` (`success` / `failed`),
`output_excerpt` (tail, ≤2000 chars). Overlap-skipped runs write nothing. Rows older than 30 days are
pruned opportunistically (~1 run in 500), so the table stays bounded without its own schedule entry.

Caveat: for `runInBackground()` tasks, `duration_ms` includes the few seconds until Laravel's
`schedule:finish` process fires the success/failure callback — treat it as "roughly how long", not a
precise timing.

Health queries (per deployment):

```sql
-- last run of every command
SELECT DISTINCT ON (command) command, status, duration_ms, finished_at
FROM scheduled_task_runs ORDER BY command, started_at DESC;

-- failures in the last 24h
SELECT command, finished_at, output_excerpt
FROM scheduled_task_runs
WHERE status = 'failed' AND started_at > now() - interval '24 hours'
ORDER BY started_at DESC;

-- a command that should have run in the last hour but didn't
SELECT 'attendance:auto-checkout' AS command
WHERE NOT EXISTS (
  SELECT 1 FROM scheduled_task_runs
  WHERE command = 'attendance:auto-checkout' AND started_at > now() - interval '5 minutes'
);
```

`schedule:run` failures are also in `/var/log/laravel/*_schedule.log` (§2) and, for every command,
`Log::error("Scheduled command failed: …")` in `storage/logs/laravel.log`.

---

## 4. Supervisor — one worker group per deployment

The DB queue is per-database, so each client's worker only ever sees that client's jobs — no
cross-client starvation, keep one group each. `/etc/supervisor/conf.d/api_gsc_worker.conf` (repeat
per client, changing the name/path/logfile):

```ini
[program:api_gsc_worker]
process_name=%(program_name)s_%(process_num)02d
command=php /var/www/html/api_gsc_backend/artisan queue:work database --sleep=3 --tries=3 --backoff=10 --max-time=3600 --max-jobs=1000 --timeout=120
directory=/var/www/html/api_gsc_backend
autostart=true
autorestart=true
stopasgroup=true
killasgroup=true
user=www-data
numprocs=1
redirect_stderr=true
stdout_logfile=/var/log/laravel/api_gsc_worker.log
stdout_logfile_maxbytes=50MB
stdout_logfile_backups=5
stopwaitsecs=130
```

Parameter notes:

- `--timeout=120` must stay **above the longest queued job**. Audit with the monitoring table plus
  `SELECT max(...)` on your job runtimes. `backup:database-auto` shells out to `pg_dump` with a
  600 s cap, but that runs in the **scheduler**, not the queue, so it does not constrain
  `--timeout`. If any real queued job (batch email/SMS, webhook retry, TikTok fetch) can exceed
  120 s, raise `--timeout` **and** the DB queue `retry_after` (`config/queue.php` →
  `connections.database.retry_after`, currently 90) so `retry_after > timeout`; otherwise a slow job
  is retried while still running.
- `--max-time=3600` + `--max-jobs=1000` recycle the worker to cap memory growth. Confirm steady-state
  RSS with `ps -o rss,cmd -p $(pgrep -f "api_gsc_backend/artisan queue:work")` over a day before
  changing these.
- `stopwaitsecs` > `--timeout` so a graceful deploy restart never `SIGKILL`s a job mid-flight.
- After `jwt:generate-secrets` runs (weekly, or manually), the command issues `queue:restart`;
  Supervisor `autorestart` brings the worker straight back with the new secret.

Apply:

```bash
sudo supervisorctl reread && sudo supervisorctl update && sudo supervisorctl status
```

All `api_*_worker:*` entries must be `RUNNING`. Spot-check logs:

```bash
sudo supervisorctl tail -100 api_gsc_worker:api_gsc_worker_00
```

### 4a. Payment commands now fan out to the queue

`students:check-payment-status` and `students:send-payment-reminders` used to evaluate every student
inline (`PaymentController::getStudentPaymentDetails()` per student) — minutes of work on the
scheduler tick for a large client. They now:

- run a cheap `COUNT` + `chunkById()` over active/suspended students and dispatch one job per chunk
  (`EvaluateStudentPaymentStatusJob` / `SendPaymentReminderBatchJob`, default 200 students/chunk,
  override with `--chunk=`), returning in ~1–2 s;
- the per-student business logic is **unchanged**, just moved into the job's `process()`;
- `--dry-run` (and `--force` for reminders) still run inline so an operator gets a summary.

Queue impact per deployment: at 00:00 a burst of `ceil(students/200)` evaluation jobs, at 08:00 a
burst of reminder jobs. Each job runs ~5–30 s for 100–200 students. With one worker per client
(§4) they drain sequentially in a few minutes — fine at 00:00/08:00. If a large client needs it
faster, raise `numprocs` for that client's worker **or** lower `--chunk`. These jobs are
`tries=1, timeout=600`; keep the worker `--timeout` ≥ the chunk runtime (200-student chunk × your
slowest client ≈ under 120 s in testing, but measure). Reminder sends themselves already go through
`SendBatchEmailJob` / `SendBatchSmsJob`, so the batch job only queues more work — it does not send
synchronously.

---

## 5. PHP-FPM sizing (per pool)

Current pool: `pm.max_children=5`, `start_servers=2`, `min_spare=1`, `max_spare=3`. Do not raise
`max_children` blind — more FPM workers against one Postgres can make latency worse. Size it from
measurement.

```bash
free -h ; nproc ; uptime
vmstat 1 5
# average RSS of a warm FPM worker:
ps -ylC php-fpm8.3 --sort:rss | awk 'NR>1 {sum+=$8; n++} END {printf "avg RSS: %.0f MB over %d procs\n", sum/n/1024, n}'
```

```
                RAM available to PHP-FPM (total − Postgres − Redis − workers − OS − headroom)
max_children ≈ ───────────────────────────────────────────────────────────────────────────────
                                  average FPM worker RSS
```

Leave headroom for: PostgreSQL (`shared_buffers` + per-connection work_mem × connections), the eight
queue workers (~§4 RSS each), Supervisor, cron PHP processes (one `schedule:run` per client per
minute, plus `runInBackground` children), the OS page cache. Across eight pools the sum of every
pool's `max_children` is what has to fit, not one pool's.

Enable the slow log to find the real fix before adding workers — `www.conf`:

```ini
slowlog = /var/log/php8.3-fpm/api_gsc.slow.log
request_slowlog_timeout = 5s
pm.status_path = /fpm-status
```

Investigate what shows up (N+1 ORM queries, external API calls in the request path, large report
generation, synchronous mail) rather than masking it with `max_children`.

---

## 6. JWT rotation behaviour (changed)

`jwt:generate-secrets` still runs weekly (Sun 02:00) and still rewrites `JWT_SECRET` /
`JWT_REFRESH_SECRET` in `.env`. It now **also**, unless `--no-reload` is passed:

1. rebuilds the config cache (`config:cache` if config was cached, else `config:clear`) — without
   this the app keeps reading the old secret from `bootstrap/cache/config.php`;
2. runs `queue:restart` so workers reload.

Consequence, by design: every active access/refresh token becomes invalid at rotation and all users
must log in again. If weekly forced logout is unacceptable, the follow-up is a **previous-secret
grace window** — verify a token against `jwt.secret`, then fall back to `jwt.previous_secret` for
e.g. 24 h — which needs changes in `config/jwt.php`, `JwtAuthMiddleware`, and
`EntranceExamAuthMiddleware`. Until then, either accept the weekly logout or widen the interval in
`routes/console.php` (`->weeklyOn(0, '02:00')` → `->monthlyOn(1, '02:00')`).

Run `php artisan jwt:generate-secrets --no-reload` when rotating by hand mid-deploy (the deploy's own
`config:cache` / worker restart then picks it up).

---

## 7. Deploy & rollback

Per deployment (`cd /var/www/html/api_<client>`):

```bash
# 1. snapshot the two schedule files + DB
cp routes/console.php routes/console.php.bak
cp bootstrap/app.php  bootstrap/app.php.bak
php artisan backup:database-auto        # or your normal pg_dump

# 2. pull the new code, install
git pull            # (or your deploy mechanism)
composer install --no-dev --optimize-autoloader

# 3. migrate (adds scheduled_task_runs only — purely additive)
php artisan migrate --force

# 4. rebuild caches
php artisan optimize:clear
php artisan config:cache && php artisan route:cache && php artisan event:cache

# 5. verify
php artisan schedule:list                       # 15 entries, each once
grep -n "command(" bootstrap/app.php             # no matches

# 6. restart workers so they load the new code
sudo supervisorctl restart api_<client>_worker:*
```

Then watch for a few cycles:

```bash
tail -f /var/log/laravel/api_<client>_schedule.log
# and, after ~2 min:
psql -d <client_db> -c "SELECT DISTINCT ON (command) command, status, finished_at
                        FROM scheduled_task_runs ORDER BY command, started_at DESC;"
```

### Rollback

```bash
mv routes/console.php.bak routes/console.php
mv bootstrap/app.php.bak  bootstrap/app.php
php artisan migrate:rollback --step=1      # drops scheduled_task_runs
php artisan optimize:clear && php artisan config:cache
sudo supervisorctl restart api_<client>_worker:*
```

The migration is additive and the two code files are self-contained, so rollback is a file swap plus
one `migrate:rollback`. No data migration, no schema change to existing tables.

---

## 8. Regression checklist (run once per deployment after rollout)

| Area | Check |
|---|---|
| Scheduler | `schedule:list` shows 15 commands, each once, at the §1 frequencies |
| Scheduler | no duplicate cron / systemd timer / in-script `schedule:run` (§2) |
| Monitoring | `scheduled_task_runs` gets `success` rows within 2 min; force one failure → `failed` row with `output_excerpt` |
| Attendance | `php artisan attendance:auto-checkout` twice → no double checkout |
| HR | `php artisan hr:mark-absent-staff` twice → 2nd run exits 0, no duplicate `(staff_id, attendance_date)` |
| Commissions | `php artisan commissions:close-expired` twice → 2nd run closes 0 |
| Communications | `php artisan communications:process-scheduled` twice → no double dispatch |
| Facebook / TikTok | manually run each of the 5 webhook/lead commands once → completes |
| Payments (dry-run) | `students:check-payment-status --dry-run`, `students:send-payment-reminders --dry-run` → summary matches pre-change counts |
| Payments (queued) | `students:check-payment-status` (no flags) returns in ~1–2 s; `jobs` table gets `EvaluateStudentPaymentStatusJob` rows; a worker drains them with 0 `failed_jobs`; student statuses end up as the dry-run predicted |
| Backups | `backup:database-auto` → file on the configured disk + `database_backups` row |
| JWT | `jwt:generate-secrets` (staging) → new token works immediately, old token rejected, no manual deploy |
| Queue | `supervisorctl status` all `RUNNING`; failed_jobs not growing abnormally |
| API | authentication, a cached endpoint, and a Facebook/TikTok webhook callback all still work |
