monitoring.metrics_storage module

Lightweight DB-backed metrics storage with safe, best-effort batching.

Design goals: - Fail-open: never break app flow if DB unavailable or pymongo missing - No import-time DB connections; initialize lazily on first flush - Use environment variables only (avoid importing config to prevent cycles) - Memory-safety in misconfiguration: drop/cap buffer when storage unavailable

Environment variables: - METRICS_DB_ENABLED: ”true/1/yes“ to enable DB writes (default: false) - MONGODB_URL: Mongo connection string (required when enabled) - DATABASE_NAME: Database name (default: code_keeper_bot) - METRICS_COLLECTION: Collection name (default: service_metrics) - METRICS_BATCH_SIZE: Batch size threshold (default: 50) - METRICS_FLUSH_INTERVAL_SEC: Time-based flush threshold (default: 5 seconds) - METRICS_MAX_BUFFER: Max queued items in memory (default: 5000) - METRICS_ROLLUP_SECONDS: Rollup bucket size in seconds for DB writes (default: 60) - METRICS_TTL_DAYS: Retention window for the collection (default: 30)

monitoring.metrics_storage.emit_event(event, severity='info', **fields)[מקור]
Return type:

None

פרמטרים:
monitoring.metrics_storage.METRICS_TTL_DAYS_DEFAULT = 30

ברירת המחדל של חלון השמירה של service_metrics, בימים.

30 ולא פחות, כי זה הטווח שהמערכת עצמה מבקשת: הדשבורד (webapp/templates/admin_observability.html) מציע כפתור 30d, ו- OBSERVABILITY_WARMUP_RANGES שואל 30 יום אחורה בכל עלייה של התהליך. חלון קצר יותר היה מרוקן את שתי התצוגות האלה בשקט — הן היו מחזירות אפס בלי שום שגיאה. ‏(אישיו #3331: ה-endpoint של התחזוקה הגדיר 24 שעות, בסתירה ישירה.)

monitoring.metrics_storage.METRICS_TTL_INDEX_NAME = 'metrics_ttl'

יצירת האינדקסים בעלייה ו-endpoint התחזוקה נוגעים באותו אינדקס, ושני שמות שונים על אותו מפתח הם IndexOptionsConflict — כלומר אחד מהם פשוט לא יחול. השם נלקח מזה שכבר קיים בפרודקשן.

Type:

שם אינדקס ה-TTL. שם אחד לכל המערכת

monitoring.metrics_storage.metrics_collection_name()[מקור]

שם האוסף שאליו הכותב באמת כותב.

מקור אמת יחיד: ה-endpointים של התחזוקה ויצירת האינדקסים חייבים לגעת באותו אוסף שהכותב כותב אליו. שם קשיח אצלם היה מנקה אוסף אחר בזמן שהאמיתי מתנפח.

Return type:

str

monitoring.metrics_storage.metrics_ttl_seconds()[מקור]

חלון השמירה של האוסף בשניות, לפי METRICS_TTL_DAYS.

נקרא בזמן השימוש ולא בזמן הייבוא, כדי שערך שמוגדר מאוחר יותר ייקרא בפועל.

Return type:

int

monitoring.metrics_storage.stale_ttl_indexes(coll, *, keep_name)[מקור]

כל אינדקס TTL באוסף שאינו זה שהמערכת מחזיקה — כלומר שריד שיש להפיל.

החוזה: לאוסף הזה יש חלון שמירה אחד, ולכן אינדקס TTL אחד — keep_name על {ts: 1}. כל אינדקס TTL אחר הוא שריד משתי גרסאות קודמות של endpoint התחזוקה, ושניהם עולים בכל כתיבה:

  • על {ts: 1} בשם אחר (ttl_cleanup_ts) — חוסם: מונגו מחזיר IndexOptionsConflict כששני אינדקסים חולקים מפתח ונבדלים באופציות, והכשל חוזר כשדה status בתשובה ולא כחריגה. כלומר שקט למי שלא קורא.

  • על שדה אחר (ttl_cleanup על timestamp) — אינרטי: אין לשדה הזה שום כותב באוסף, ולכן ”If a document does not contain the indexed field, the document will not expire“ (מקור: https://www.mongodb.com/docs/manual/core/index-ttl/). הוא לא חוסם את היצירה ולכן לא מרגישים בו — הוא רק תופס מקום ומתוחזק בכל כתיבה, לנצח.

הפונקציה מחזירה את שניהם. זהו החוזה ולא רשימת מופעים: אינדקס TTL חדש על האוסף הזה מתחיל מהמודול הזה, לא מהצד.

Return type:

List[str]

פרמטרים:
monitoring.metrics_storage.flush(force=False)[מקור]
Return type:

None

פרמטרים:

force (bool)

monitoring.metrics_storage.enqueue_request_metric(status_code, duration_seconds, *, request_id=None, extra=None)[מקור]

Queue a single request metric for best-effort DB persistence.

This write path is opt-in via METRICS_DB_ENABLED=true.

In production we do not write ”one document per request“. Instead, we roll up requests into per-bucket documents to keep MongoDB load low.

Return type:

None

פרמטרים:
monitoring.metrics_storage.aggregate_request_timeseries(*, start_dt, end_dt, granularity_seconds)[מקור]

Aggregate request metrics into fixed time buckets.

Return type:

List[Dict[str, Any]]

פרמטרים:
monitoring.metrics_storage.aggregate_top_endpoints(*, start_dt, end_dt, limit=5)[מקור]

Return the slowest HTTP endpoints within the given time window.

Return type:

List[Dict[str, Any]]

פרמטרים:
monitoring.metrics_storage.average_request_duration(*, start_dt, end_dt)[מקור]

Return the average request duration for a given window.

Return type:

Optional[float]

פרמטרים:
monitoring.metrics_storage.aggregate_error_ratio(*, start_dt, end_dt)[מקור]

Return total/error counts for the window.

Return type:

Dict[str, int]

פרמטרים:
monitoring.metrics_storage.find_by_request_id(request_id, *, limit=20)[מקור]

Find metrics records by request_id.

Used by triage service to provide fallback when Sentry is unavailable. Returns empty list on any failure (fail-open).

Return type:

List[Dict[str, Any]]

פרמטרים:
  • request_id (str)

  • limit (int)

monitoring.metrics_storage.aggregate_latency_percentiles(*, start_dt, end_dt, percentiles=(50, 95, 99), sample_limit=5000)[מקור]

Return latency percentiles (seconds) for the given window.

Best-effort: - Try Mongo $percentile aggregation when available. - Otherwise, sample up to sample_limit records and compute percentiles in Python.

Return type:

Dict[str, float]

פרמטרים: