280 lines
12 KiB
Markdown
280 lines
12 KiB
Markdown
# HTTPX Logging Handler Example
|
|
|
|
Use this reference when a Python application needs to send log records to an HTTP endpoint with [`httpx`](https://www.python-httpx.org/). The example follows the same boundaries as the TCP network example: feature modules use normal named loggers, application startup configures logging once, a queue keeps HTTP I/O off the caller path, and the receiver or collector owns final routing.
|
|
|
|
Source docs to keep nearby:
|
|
|
|
- [HTTPX clients](https://www.python-httpx.org/advanced/clients/) for connection pooling and shared request configuration.
|
|
- [HTTPX timeouts](https://www.python-httpx.org/advanced/timeouts/) for connect, read, write, and pool timeout behavior.
|
|
- [HTTPX JSON requests](https://www.python-httpx.org/quickstart/#sending-json-encoded-data) for posting JSON payloads.
|
|
- [HTTPX exceptions](https://www.python-httpx.org/quickstart/#exceptions) for `RequestError`, `HTTPStatusError`, and `HTTPError` handling.
|
|
- [Dealing with handlers that block](https://docs.python.org/3/howto/logging-cookbook.html#dealing-with-handlers-that-block) for why slow handlers should sit behind `QueueHandler` and `QueueListener`.
|
|
- [`QueueHandler`](https://docs.python.org/3/library/logging.handlers.html#queuehandler) and [`QueueListener`](https://docs.python.org/3/library/logging.handlers.html#queuelistener) for queue-backed logging mechanics.
|
|
- [`logging.makeLogRecord`](https://docs.python.org/3/library/logging.html#logging.makeLogRecord) for rebuilding records from serialized fields.
|
|
|
|
## Minimal Topology
|
|
|
|
Send each record as a JSON HTTP request to a collector endpoint. Use `httpx.Client`, not the top-level `httpx.post`, because a handler may send many records to the same host and should reuse connections.
|
|
|
|
```text
|
|
application code -> named logger -> QueueHandler -> QueueListener -> HTTPX JSON handler -> HTTP receiver or collector
|
|
```
|
|
|
|
Application modules stay ordinary:
|
|
|
|
```python title="feature.py"
|
|
import logging
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def sync_customer(customer_id: str) -> None:
|
|
logger.info("Syncing customer %s", customer_id)
|
|
```
|
|
|
|
The module does not know whether logs are written locally, sent over HTTP, or forwarded by a platform collector.
|
|
|
|
## Minimal Receiver
|
|
|
|
This receiver is for local testing. A production deployment would usually send the same JSON shape to a managed log collector, OpenTelemetry collector, service endpoint, or internal ingestion API.
|
|
|
|
```python title="log_http_receiver.py"
|
|
import json
|
|
import logging
|
|
import logging.config
|
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
|
|
|
LOGGING = {
|
|
"version": 1,
|
|
"disable_existing_loggers": False,
|
|
"formatters": {
|
|
"console": {
|
|
"format": "%(asctime)s %(levelname)s %(name)s %(message)s",
|
|
"datefmt": "%Y-%m-%d %H:%M:%S",
|
|
}
|
|
},
|
|
"handlers": {
|
|
"console": {
|
|
"class": "logging.StreamHandler",
|
|
"formatter": "console",
|
|
"stream": "ext://sys.stdout",
|
|
}
|
|
},
|
|
"root": {"level": "INFO", "handlers": ["console"]},
|
|
}
|
|
|
|
|
|
class LogRecordRequestHandler(BaseHTTPRequestHandler):
|
|
def do_POST(self) -> None:
|
|
if self.path != "/logs":
|
|
self.send_error(404)
|
|
return
|
|
|
|
length = int(self.headers.get("Content-Length", "0"))
|
|
body = self.rfile.read(length)
|
|
|
|
try:
|
|
payload = json.loads(body.decode("utf-8"))
|
|
record = logging.makeLogRecord(payload)
|
|
except Exception:
|
|
logging.getLogger(__name__).exception("Dropped malformed log record")
|
|
self.send_error(400)
|
|
return
|
|
|
|
logger = logging.getLogger(record.name)
|
|
if logger.isEnabledFor(record.levelno):
|
|
logger.handle(record)
|
|
|
|
self.send_response(204)
|
|
self.end_headers()
|
|
|
|
def log_message(self, format: str, *args: object) -> None:
|
|
logging.getLogger("http.server").debug(format, *args)
|
|
|
|
|
|
def main() -> None:
|
|
logging.config.dictConfig(LOGGING)
|
|
server = ThreadingHTTPServer(("127.0.0.1", 9021), LogRecordRequestHandler)
|
|
with server:
|
|
logging.getLogger(__name__).info("Listening for HTTP log records")
|
|
server.serve_forever()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|
|
```
|
|
|
|
### Receiver Mechanics
|
|
|
|
- The receiver accepts `POST /logs` with a JSON body and returns `204` when the record is accepted.
|
|
- `makeLogRecord` turns the JSON dictionary back into a standard `LogRecord`, so the receiver can use the normal logger hierarchy.
|
|
- Receiver-side filtering still works, but client-side filtering is better when volume matters because it avoids serializing and transmitting records that will be discarded.
|
|
- `log_message` is redirected into the logging system at `DEBUG` so access logs do not pollute normal output.
|
|
- Binding to `127.0.0.1` keeps the demo local. If the receiver is reachable across a network, put authentication, TLS, rate limits, and request size limits in front of it.
|
|
|
|
## Client Configuration
|
|
|
|
The client side uses a queue-backed logging handler. The listener thread owns the `httpx.Client`, posts JSON records, and closes the connection pool during shutdown.
|
|
|
|
```python title="logging_config.py"
|
|
import copy
|
|
import logging
|
|
import logging.handlers
|
|
import queue
|
|
|
|
import httpx
|
|
|
|
|
|
class PreservingQueueHandler(logging.handlers.QueueHandler):
|
|
def prepare(self, record: logging.LogRecord) -> logging.LogRecord:
|
|
copied = copy.copy(record)
|
|
copied.message = copied.getMessage()
|
|
copied.msg = copied.message
|
|
copied.args = None
|
|
|
|
if copied.exc_info is not None and copied.exc_text is None:
|
|
copied.exc_text = logging.Formatter().formatException(copied.exc_info)
|
|
copied.exc_info = None
|
|
|
|
return copied
|
|
|
|
|
|
class HttpxJsonLogHandler(logging.Handler):
|
|
def __init__(self, collector_url: str, token: str | None = None) -> None:
|
|
super().__init__()
|
|
headers = {"content-type": "application/json"}
|
|
if token is not None:
|
|
headers["authorization"] = f"Bearer {token}"
|
|
|
|
timeout = httpx.Timeout(2.0, connect=1.0, write=2.0, pool=1.0)
|
|
self.client = httpx.Client(base_url=collector_url, headers=headers, timeout=timeout)
|
|
|
|
def emit(self, record: logging.LogRecord) -> None:
|
|
try:
|
|
response = self.client.post("/logs", json=self._payload_from_record(record))
|
|
response.raise_for_status()
|
|
except httpx.HTTPError:
|
|
self.handleError(record)
|
|
|
|
def close(self) -> None:
|
|
self.client.close()
|
|
super().close()
|
|
|
|
def _payload_from_record(self, record: logging.LogRecord) -> dict[str, object]:
|
|
payload: dict[str, object] = {
|
|
"name": record.name,
|
|
"levelno": record.levelno,
|
|
"levelname": record.levelname,
|
|
"pathname": record.pathname,
|
|
"lineno": record.lineno,
|
|
"funcName": record.funcName,
|
|
"created": record.created,
|
|
"process": record.process,
|
|
"processName": record.processName,
|
|
"threadName": record.threadName,
|
|
"msg": record.getMessage(),
|
|
"args": None,
|
|
}
|
|
if record.exc_text is not None:
|
|
payload["exc_text"] = record.exc_text
|
|
return payload
|
|
|
|
|
|
def configure_logging(
|
|
collector_url: str = "http://127.0.0.1:9021",
|
|
token: str | None = None,
|
|
) -> logging.handlers.QueueListener:
|
|
log_queue: queue.Queue[logging.LogRecord] = queue.Queue(maxsize=1000)
|
|
queue_handler = PreservingQueueHandler(log_queue)
|
|
http_handler = HttpxJsonLogHandler(collector_url, token)
|
|
listener = logging.handlers.QueueListener(log_queue, http_handler, respect_handler_level=True)
|
|
|
|
root = logging.getLogger()
|
|
root.setLevel(logging.INFO)
|
|
root.handlers[:] = [queue_handler]
|
|
|
|
listener.start()
|
|
return listener
|
|
```
|
|
|
|
```python title="app.py"
|
|
import logging
|
|
|
|
from logging_config import configure_logging
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def main() -> None:
|
|
listener = configure_logging()
|
|
try:
|
|
logger.info("Service started")
|
|
logger.warning("Example warning from the HTTPX client")
|
|
finally:
|
|
listener.stop()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|
|
```
|
|
|
|
Start `log_http_receiver.py` first, then run `app.py`. The receiver should print the records using its own formatter.
|
|
|
|
## HTTPX Mechanics
|
|
|
|
- `httpx.Client` keeps a connection pool. That matters for log handlers because repeated top-level `httpx.post(...)` calls would create new connections instead of reusing them.
|
|
- `base_url` makes the handler explicit about the collector host while keeping the endpoint path short.
|
|
- `timeout` is explicit. HTTPX has default timeouts, but logging code should state its tolerance for connect, write, read, and pool waits.
|
|
- `response.raise_for_status()` turns non-2xx responses into `HTTPStatusError`, which then goes through the logging handler's normal error path.
|
|
- `close()` closes the HTTPX connection pool. Pair this with `listener.stop()` during application shutdown so queued records are sent and resources are released.
|
|
|
|
## Logging Mechanics
|
|
|
|
- The application still configures logging once. Feature modules only call named loggers.
|
|
- The HTTP handler sits behind `QueueHandler` and `QueueListener` because HTTP requests can block on DNS, connection pooling, TLS, request writes, response reads, and collector back pressure.
|
|
- `PreservingQueueHandler` copies the record and formats exception text before clearing `exc_info`, so the queued record is safe to serialize and still carries useful traceback information.
|
|
- The handler serializes a deliberate subset of `LogRecord` fields. Sending `record.__dict__` wholesale is easy, but it can include unserializable objects, accidental high-cardinality fields, or data the collector should not receive.
|
|
- Authentication is represented as an optional bearer token header. In real applications, read tokens from a secret manager or runtime configuration, not from source code.
|
|
|
|
## Rationale Behind The Pattern
|
|
|
|
### Prefer HTTP When The Receiver Is Already An HTTP API
|
|
|
|
HTTP is a good fit when logs go to a collector, gateway, ingestion service, or internal API that already expects JSON over HTTPS. It also gives you familiar deployment controls: TLS termination, authentication, reverse proxies, rate limiting, request size limits, and conventional status codes.
|
|
|
|
### Keep HTTP Off The Caller Path
|
|
|
|
Even a fast collector can become slow during deploys, network incidents, or downstream outages. Queueing makes that failure mode a logging concern instead of a request-latency concern.
|
|
|
|
### Use A Bounded Queue For Honest Back Pressure
|
|
|
|
The example uses `queue.Queue(maxsize=1000)` so overload becomes visible. The default `QueueHandler.enqueue()` uses `put_nowait()`, so a full queue calls `handleError()`. For production, decide whether to drop logs, block briefly, spill to disk, or switch to a platform collector.
|
|
|
|
### Filter Before Sending
|
|
|
|
The collector can reject records, but rejected records already consumed CPU, queue capacity, and network bandwidth. Use client-side logger and handler levels to avoid sending noisy records unless the deployment explicitly needs them.
|
|
|
|
## Production Checklist
|
|
|
|
Before using this beyond a local demo:
|
|
|
|
1. Use HTTPS and authenticate clients. Treat the collector endpoint as an ingestion boundary, not a public anonymous API.
|
|
2. Set request size limits and reject malformed payloads early.
|
|
3. Decide the outage policy for collector failures and full queues.
|
|
4. Add service, environment, instance, request, trace, or tenant identifiers as explicit serialized fields when operators need correlation.
|
|
5. Redact or avoid secrets before records leave the process.
|
|
6. Normalize untrusted newline-containing values if the final destination is line-oriented.
|
|
7. Tune timeouts and queue size under load, not only with a happy-path local receiver.
|
|
8. Prefer a managed collector, OpenTelemetry pipeline, or platform-native logging when one already exists.
|
|
|
|
## Review Questions
|
|
|
|
Use these questions when reviewing HTTPX logging code:
|
|
|
|
- Does application code only call named loggers, without direct HTTP calls from feature modules?
|
|
- Is the HTTP handler behind a queue for web, async, worker, or high-throughput paths?
|
|
- Does the handler use a reusable `httpx.Client` rather than top-level request functions?
|
|
- Are timeouts explicit and short enough for a logging path?
|
|
- Are collector failures, non-2xx responses, and full queues handled deliberately?
|
|
- Does shutdown stop the listener and close the HTTPX client?
|
|
- Are authentication, TLS, secrets, request size, and high-cardinality context handled deliberately? |