monitoring.metrics_storage module
Lightweight DB-backed metrics storage with safe, best-effort batching.
Design goals: - Fail-open: never break app flow if DB unavailable or pymongo missing - No import-time DB connections; initialize lazily on first flush - Use environment variables only (avoid importing config to prevent cycles) - Memory-safety in misconfiguration: drop/cap buffer when storage unavailable
Environment variables: - METRICS_DB_ENABLED: ”true/1/yes“ to enable DB writes (default: false) - MONGODB_URL: Mongo connection string (required when enabled) - DATABASE_NAME: Database name (default: code_keeper_bot) - METRICS_COLLECTION: Collection name (default: service_metrics) - METRICS_BATCH_SIZE: Batch size threshold (default: 50) - METRICS_FLUSH_INTERVAL_SEC: Time-based flush threshold (default: 5 seconds) - METRICS_MAX_BUFFER: Max queued items in memory (default: 5000) - METRICS_ROLLUP_SECONDS: Rollup bucket size in seconds for DB writes (default: 60) - METRICS_TTL_DAYS: Retention window for the collection (default: 30)
- monitoring.metrics_storage.METRICS_TTL_DAYS_DEFAULT = 30
ברירת המחדל של חלון השמירה של
service_metrics, בימים.30 ולא פחות, כי זה הטווח שהמערכת עצמה מבקשת: הדשבורד (
webapp/templates/admin_observability.html) מציע כפתור30d, ו-OBSERVABILITY_WARMUP_RANGESשואל 30 יום אחורה בכל עלייה של התהליך. חלון קצר יותר היה מרוקן את שתי התצוגות האלה בשקט — הן היו מחזירות אפס בלי שום שגיאה. (אישיו #3331: ה-endpoint של התחזוקה הגדיר 24 שעות, בסתירה ישירה.)
- monitoring.metrics_storage.METRICS_TTL_INDEX_NAME = 'metrics_ttl'
יצירת האינדקסים בעלייה ו-endpoint התחזוקה נוגעים באותו אינדקס, ושני שמות שונים על אותו מפתח הם IndexOptionsConflict — כלומר אחד מהם פשוט לא יחול. השם נלקח מזה שכבר קיים בפרודקשן.
- Type:
שם אינדקס ה-TTL. שם אחד לכל המערכת
- monitoring.metrics_storage.metrics_collection_name()[מקור]
שם האוסף שאליו הכותב באמת כותב.
מקור אמת יחיד: ה-endpointים של התחזוקה ויצירת האינדקסים חייבים לגעת באותו אוסף שהכותב כותב אליו. שם קשיח אצלם היה מנקה אוסף אחר בזמן שהאמיתי מתנפח.
- Return type:
- monitoring.metrics_storage.metrics_ttl_seconds()[מקור]
חלון השמירה של האוסף בשניות, לפי
METRICS_TTL_DAYS.נקרא בזמן השימוש ולא בזמן הייבוא, כדי שערך שמוגדר מאוחר יותר ייקרא בפועל.
- Return type:
- monitoring.metrics_storage.stale_ttl_indexes(coll, *, keep_name)[מקור]
כל אינדקס TTL באוסף שאינו זה שהמערכת מחזיקה — כלומר שריד שיש להפיל.
החוזה: לאוסף הזה יש חלון שמירה אחד, ולכן אינדקס TTL אחד —
keep_nameעל{ts: 1}. כל אינדקס TTL אחר הוא שריד משתי גרסאות קודמות של endpoint התחזוקה, ושניהם עולים בכל כתיבה:על
{ts: 1}בשם אחר (ttl_cleanup_ts) — חוסם: מונגו מחזירIndexOptionsConflictכששני אינדקסים חולקים מפתח ונבדלים באופציות, והכשל חוזר כשדהstatusבתשובה ולא כחריגה. כלומר שקט למי שלא קורא.על שדה אחר (
ttl_cleanupעלtimestamp) — אינרטי: אין לשדה הזה שום כותב באוסף, ולכן ”If a document does not contain the indexed field, the document will not expire“ (מקור: https://www.mongodb.com/docs/manual/core/index-ttl/). הוא לא חוסם את היצירה ולכן לא מרגישים בו — הוא רק תופס מקום ומתוחזק בכל כתיבה, לנצח.
הפונקציה מחזירה את שניהם. זהו החוזה ולא רשימת מופעים: אינדקס TTL חדש על האוסף הזה מתחיל מהמודול הזה, לא מהצד.
- monitoring.metrics_storage.enqueue_request_metric(status_code, duration_seconds, *, request_id=None, extra=None)[מקור]
Queue a single request metric for best-effort DB persistence.
This write path is opt-in via
METRICS_DB_ENABLED=true.In production we do not write ”one document per request“. Instead, we roll up requests into per-bucket documents to keep MongoDB load low.
- monitoring.metrics_storage.aggregate_request_timeseries(*, start_dt, end_dt, granularity_seconds)[מקור]
Aggregate request metrics into fixed time buckets.
- monitoring.metrics_storage.aggregate_top_endpoints(*, start_dt, end_dt, limit=5)[מקור]
Return the slowest HTTP endpoints within the given time window.
- monitoring.metrics_storage.average_request_duration(*, start_dt, end_dt)[מקור]
Return the average request duration for a given window.
- monitoring.metrics_storage.aggregate_error_ratio(*, start_dt, end_dt)[מקור]
Return total/error counts for the window.
- monitoring.metrics_storage.find_by_request_id(request_id, *, limit=20)[מקור]
Find metrics records by request_id.
Used by triage service to provide fallback when Sentry is unavailable. Returns empty list on any failure (fail-open).
- monitoring.metrics_storage.aggregate_latency_percentiles(*, start_dt, end_dt, percentiles=(50, 95, 99), sample_limit=5000)[מקור]
Return latency percentiles (seconds) for the given window.
Best-effort: - Try Mongo $percentile aggregation when available. - Otherwise, sample up to sample_limit records and compute percentiles in Python.