Key Concepts
- Prometheus: A tool for monitoring applications and extracting metrics.
- Grafana: A tool for visualizing metrics obtained from Prometheus.
- Flask: A Python web framework used to create the example application.
- Docker & Docker Compose: Containerization tools used to deploy the application, Prometheus, and Grafana.
- Metrics: Numerical data points that represent the state or behavior of an application.
- Counter: A Prometheus metric type that represents a cumulative value that only increases.
- Summary: A Prometheus metric type that aggregates data over time, providing statistics like averages and quantiles.
- Endpoints: Specific URLs in a web application that handle requests.
- Error Handling: Mechanisms for catching and managing exceptions in an application.
- Scraping: The process by which Prometheus collects metrics from applications.
- Data Source: A connection in Grafana to a source of data, such as Prometheus.
- Dashboard: A collection of visualizations in Grafana that display metrics.
Monitoring Applications with Prometheus and Grafana
Introduction
The video demonstrates how to professionally monitor applications using Prometheus and Grafana, two open-source tools commonly used in production environments. The goal is to provide a basic setup using a Flask application, Docker, and Docker Compose, focusing on getting started rather than an in-depth tutorial.
Prometheus and Grafana Overview
- Prometheus: Collects metrics from applications.
- Grafana: Visualizes the metrics collected by Prometheus.
- The application sends metrics to Prometheus, and Grafana retrieves and displays them.
Setting up the Flask Application
- Requirements File (requirements.txt):
- Includes
Flaskandprometheus_clientpackages.
- Includes
- Flask Application (app.py):
- Imports necessary modules:
flask,request,responsefromflask, andcounter,generate_latest,CONTENT_TYPE_LATESTfromprometheus_client. - Creates a Flask app instance:
app = Flask(__name__). - Defines several endpoints:
/hello: Returns "Hello, World!"./print_number: Accepts a number as a URL parameter and returns it, also tracking the numbers for average calculation./crash: Raises aKeyErrorto simulate an application crash./metrics: Returns Prometheus metrics in the required format.
- Imports necessary modules:
- Exception Handling:
- Uses
app.errorhandlerto catch unhandled exceptions. - Creates a
Countermetric calledexceptionsto track the total number of unhandled exceptions, grouped by endpoint and exception type. - The
catch_allfunction increments theexceptionscounter each time an unhandled exception occurs.
- Uses
- Request Counting:
- Uses
app.before_requestto count all incoming requests. - Creates a
Countermetric calledrequest_countto track the total number of requests, grouped by method (GET, POST, etc.) and endpoint. - The
count_requestsfunction increments therequest_countcounter before each request.
- Uses
- Number Summarization:
- Creates a
Summarymetric calledprint_numberto track the numbers passed to the/print_numberendpoint. - The
observemethod of theSummaryobject is used to record each number.
- Creates a
- Metrics Endpoint:
- The
/metricsendpoint returns a response object containing the generated metrics data usinggenerate_latest(). - The
CONTENT_TYPE_LATESTis used to set the correct MIME type for the response.
- The
Docker and Docker Compose Setup
- Dockerfile:
- Uses
python:3.12-slimas the base image. - Sets the working directory to
/app. - Copies
requirements.txtand installs the dependencies usingpip install -r requirements.txt. - Copies
app.py. - Exposes port 5000.
- Sets the command to run the application:
python app.py.
- Uses
- Prometheus Configuration (prometheus.yaml):
- Sets the
scrape_intervalto 5 seconds. - Configures a
job_namecalled "flask_app". - Specifies the target as
flask:5000, where "flask" is the hostname of the Flask application container.
- Sets the
- Docker Compose File (docker-compose.yaml):
- Defines three services:
flask,prometheus, andgrafana. - Flask:
- Builds the application using the Dockerfile in the current directory.
- Sets the container name to "flask_app".
- Maps port 5000 to port 5000.
- Prometheus:
- Uses the
prom/prometheus:latestimage. - Sets the container name to "prometheus".
- Mounts the
prometheus.yamlfile to/etc/prometheus/prometheus.yamlas read-only. - Maps port 9090 to port 9090.
- Uses the
- Grafana:
- Uses the
grafana/grafana:latestimage. - Sets the container name to "grafana".
- Maps port 3000 to port 3000.
- Specifies that it
depends_onPrometheus.
- Uses the
- Defines three services:
Running the Application
- Build the Docker images using
docker compose build. - Start the containers using
docker compose up. - Access the Flask application at
localhost:5000. - Access Prometheus at
localhost:9090. - Access Grafana at
localhost:3000.
Configuring Grafana
- Log in to Grafana with the default credentials (admin/admin).
- Add a data source:
- Select Prometheus as the data source.
- Set the URL to
http://prometheus:9090. - Save and test the connection.
- Create a new dashboard and add visualizations:
- Select Prometheus as the data source.
- Choose a metric to visualize, such as
app_requests_totalorexceptions_total. - Use the code editor to aggregate metrics using functions like
sum. - Save the dashboard.
Examples and Demonstrations
- The video demonstrates how to visualize the total number of requests over time using the
app_requests_totalmetric. - It shows how to visualize the total number of exceptions, grouped by exception type, using the
exceptions_totalmetric. - It explains how to calculate the average of the numbers passed to the
/print_numberendpoint by dividing theapp_print_number_summary_summetric by theapp_print_number_summary_countmetric. - The video also shows how to trigger exceptions by accessing the
/crashendpoint and passing non-numeric values to the/print_numberendpoint.
Conclusion
The video provides a practical guide to setting up a basic monitoring system using Prometheus and Grafana with a Flask application and Docker Compose. It covers the essential steps for collecting, visualizing, and analyzing application metrics, enabling professional monitoring and troubleshooting. The presenter encourages viewers to explore further customization and more advanced features of Prometheus and Grafana.
AI summaries can miss context or contain errors. Check important details against the original video.





