Monitoring¶
For the person running a Kosmos server. This page covers the health check URLs, where each log is written, the error emails the application sends, how to see failed background tasks, and what to keep an eye on.
Kosmos ships three health check URLs, writes logs and installs log rotation. It does not ship dashboards, metrics or alerting. Where something is missing, this page says so and gives the usual way to fill the gap.
Health checks¶
All three URLs answer without a login, accept only GET and HEAD, return a
small JSON body and ask not to be cached. They are defined in
config/health.py.
| URL | What it checks | Healthy | Unhealthy |
|---|---|---|---|
/health/live/ |
That the web application answers. It does not touch the database. | 200 {"status": "ok"} |
No response, or an error from nginx (502, 504). |
/health/ready/ |
That the web application can run a query (SELECT 1) against the database. |
200 {"status": "ok"} |
503 {"status": "unavailable"} |
/health/worker/ |
That the background worker is processing its schedules: no scheduled job has been left overdue for more than five minutes. | 200 {"status": "ok"} |
503 {"status": "unavailable"}, also when setup_schedules has never been run. |
Point an uptime monitor at /health/ready/ and another at
/health/worker/:
curl -fsS https://kosmos.example.com/health/ready/
curl -fsS https://kosmos.example.com/health/worker/
To test the application without going through nginx, talk to its socket directly:
curl -fsS --unix-socket /run/law.sock \
-H 'Host: kosmos.example.com' http://localhost/health/ready/
Two limits to know about:
- The request must carry a hostname listed in
ALLOWED_HOSTS. A check sent to127.0.0.1or to the server's IP address is answered with400. Use the public hostname, or send it as theHostheader as above. /health/ready/says nothing about the background worker. The web application stays healthy while the worker is stopped, which is why/health/worker/exists. It is answered by the web application, from what the worker leaves in the database, so it works while the worker is down. After a restart of the worker it can take a minute to recover. The same test drives the notice administrators see at the top of every page while the worker is down (see Background worker).
Where the logs are¶
On a production install made by scripts/install.sh --prod:
| Log | Written by | Contains |
|---|---|---|
logs/django.log in the checkout |
The web application and the worker | Warnings and errors from the application and its libraries, each with its time, level and source. Failed background tasks appear here. Start here. |
logs/error.log |
gunicorn | gunicorn's own messages (start, stop, a request ended for running past 30 seconds) and everything the web application prints, which includes a second copy of its warnings and errors. With EMAIL_BACKEND=console, outgoing email is printed here instead of being sent. |
logs/access.log |
gunicorn | One line per request that reached the application. |
journalctl -u qcluster |
The worker | The worker's warnings and errors. With EMAIL_BACKEND=console, the email the worker would send (for example the daily digest) is printed here. |
journalctl -u law |
systemd | Starts, stops and crashes of the web application. gunicorn writes its own output to logs/, so little else appears here. |
/var/log/nginx/access.log, /var/log/nginx/error.log |
nginx | Every request, and nginx's own errors, including requests refused by the rate limits and 502 responses while the application is down. These are your distribution's default locations: the Kosmos site does not set its own. |
Details that explain what you will and will not find:
- The application logs at
WARNINGand above. Routine events, such as a task finishing or a schedule firing, are not recorded. - Requests with a
Hostheader that is not inALLOWED_HOSTSare dropped from the logs on purpose. Internet scanners send these constantly. - On a development machine there is no gunicorn.
runserverandqclusterprint to their terminals, and both also writelogs/django.log.
The logging configuration is the LOGGING setting in
config/settings.py,
and gunicorn's is in your gunicorn.conf.py.
Log rotation¶
scripts/install.sh --prod installs /etc/logrotate.d/kosmos from
deploy/logrotate/kosmos.
It rotates logs/django.log, logs/error.log and logs/access.log
weekly, keeps twelve compressed copies, and uses copytruncate: the
application and gunicorn keep these files open, so a rotation that renamed
the file would leave them writing to the renamed one.
On a server that was installed by hand, copy that template to
/etc/logrotate.d/kosmos yourself and replace @APP_DIR@ with the path
of the checkout and @USER@ with the account that owns it. Without it the
three files grow until the disk is full.
The nginx logs are rotated by the logrotate configuration that your distribution's nginx package installs, and journald limits its own size.
Error emails¶
Kosmos does not change Django's standard error reporting. When DEBUG
is False, an unhandled error during a request (the user sees a "Server
Error (500)" page) is emailed to everyone in ADMINS, with the traceback
and the details of the request.
For that to reach anyone:
-
Set
ADMINSinconfig/.env. It is a Python list of name and address pairs: -
Configure outgoing email (
EMAIL_BACKEND=smtpand theEMAIL_*settings). With the installer's initialEMAIL_BACKEND=console, error reports are printed intologs/error.loginstead of being sent. -
Set
SERVER_EMAILto a sender address your mail provider accepts. Error reports are sent from it. -
Restart both services, then send a test message:
All of these variables are described in the environment variable reference.
What else the ADMINS list receives:
- Reversed online payments. When a payment processor reports that a
payment or trust deposit it had accepted has since failed or been
returned, Kosmos emails
ADMINSwith the transaction, what it changed, and a request to follow up with the client. These are business notices, so include someone who handles billing.
What is not emailed:
- Failed background tasks. A failed OCR job, sync or AI task is logged and recorded, and nobody is notified. See the next section.
- Errors that the application catches and logs itself. They are in
logs/django.logonly. - Requests ended by gunicorn's 30 second limit. These appear in
logs/error.logas a worker timeout.
An error report can quote data from the failing request, which may include client information. Send them only to mailboxes that are fit to hold it.
Failed and queued background tasks¶
The task library's own commands show the state of the background worker, from the checkout:
| Command | Shows |
|---|---|
.venv/bin/python manage.py qinfo |
A summary: queued, scheduled, successful and failed task counts, and the cluster's uptime. |
.venv/bin/python manage.py qmonitor |
The same, refreshed live, with each worker process's current task. Quit with Ctrl-C. |
For the tasks themselves, the records are in the django_q_task (finished,
with success false for failures and the error text in result),
django_q_ormq (queued) and django_q_schedule (recurring) tables. The
quickest view from the checkout:
.venv/bin/python manage.py shell -c "
from django_q.models import Failure
for t in Failure.objects.order_by('-started')[:20]:
print(t.started, t.func, str(t.result)[:200])"
A failed task is run again from the same place:
.venv/bin/python manage.py shell -c "
from django_q.tasks import async_task
from django_q.models import Failure
t = Failure.objects.get(id='<task id>')
async_task(t.func, *t.args, **(t.kwargs or {}))"
How to read them:
- A queue that keeps growing means the worker is stopped or stuck. A handful of entries that come and go is normal: the sync jobs run every minute or two.
- A task in Failed tasks may still be retrying. A failed task is attempted up to ten times, fifteen minutes apart, before it is given up on.
- Failed tasks are never cleared automatically. Delete them
(
Failure.objects.filter(...).delete()in the shell) once you have dealt with the cause. - The same failures are written to
logs/django.log, as lines containing[ERROR] django-q: Failed.
What to watch¶
Kosmos sends no alert for any of these. Check them with whatever monitoring you already run.
| Watch | How | Why |
|---|---|---|
| The site answers | An uptime monitor on /health/ready/. |
Covers nginx, gunicorn and the database together. |
| The worker is running | An uptime monitor on /health/worker/, or systemctl is-active qcluster.service. |
A stopped worker does not affect /health/ready/. Documents stop being processed and syncs stop, silently. |
| Failed tasks | manage.py qinfo, or search logs/django.log for django-q: Failed. |
The only sign that OCR, a sync or an AI job is failing. |
| Disk space: uploads | du -sh media/ with local storage. |
Every uploaded and mirrored document is stored here. |
| Disk space: logs | du -sh logs/ |
Rotated weekly by the installed logrotate file. Unbounded on a hand-built install until you add it. |
| Disk space: database | sudo -u postgres psql -c '\l+' |
Extracted document text, synced email, AI conversations and the change history of every record all live in the database. |
| TLS certificate | sudo certbot certificates |
Certbot renews automatically. Check that it is doing so. |
| Backups | That the last one finished, and that a restore works. | See Backup and restore. |
If the database's change history grows too large, clean_history can
trim it. Read its entry in Background worker before using
it.
Check that monitoring works¶
curl -fsS https://kosmos.example.com/health/ready/prints{"status": "ok"}.sendtestemail --adminsdelivers a message to theADMINSaddresses.- Stop the worker for a minute (
sudo systemctl stop qcluster.service), confirm that your worker check notices, and start it again.