Dhruvraj Singh Shekhawat

Exit 5Workshop

Driver Drowsiness Detection and Alarm System

Building a Driver Drowsiness Detection and Alarm System with Python and MediaPipe

Building a Driver Drowsiness Detection and Alarm System

A real-time, camera-based prototype that tracks eye closure, yawning and head movement with Python, OpenCV and MediaPipe, and sounds an alarm when fatigue builds up.

p1 p2 p3 p4 p5 p6
Eye open EAR = (9.0 + 9.0) / (2 × 30) = 0.30
9.0 px
Drag the slider to close a 30-pixel-wide eye and watch EAR fall. The 0.20 threshold is illustrative; the real system sets its threshold from each driver's calibration, and a single frame only measures openness. The full system waits to see how long the eye stays closed.
In this post

Driver fatigue is one of the major challenges in road safety, particularly during long journeys and low-attention driving conditions. A driver may experience prolonged eye closure, frequent yawning, reduced blinking, or involuntary head movements before reaching a dangerous level of drowsiness. Detecting these signs early can provide an opportunity to alert the driver before the situation becomes critical.

The Driver Drowsiness Detection and Alarm System is a real-time, camera-based driver monitoring prototype that uses computer vision and facial landmarks to identify potential signs of driver fatigue. The system uses MediaPipe for facial landmark detection, calculates Eye Aspect Ratio (EAR) and Mouth Aspect Ratio (MAR), estimates head pose, and analyses these measurements over time to determine a drowsiness score.

When prolonged drowsiness is detected, the system activates an alarm through a computer speaker, Raspberry Pi GPIO, Arduino, or a simulated software alarm. The project can run on a regular laptop as well as on a Raspberry Pi 4/5.

Why Use a Driver Drowsiness Detection System?

Traditional driver monitoring often relies on the driver's own awareness of fatigue. However, signs such as prolonged eye closure, yawning and head dropping can occur before the driver consciously recognises that they are becoming drowsy.

A camera-based monitoring system provides a non-contact method of continuously observing these indicators. Some of the main advantages of this approach include:

  • Real-Time Monitoring: The driver's face can be analysed continuously using a standard camera.
  • Non-Invasive Detection: The system does not require wearable sensors or physical contact with the driver.
  • Multiple Indicators: Eye closure, blinking, yawning and head movement are analysed together.
  • Personalised Calibration: The system learns the driver's normal eye and head position during startup.
  • Hardware Flexibility: The alarm can operate through a laptop, Raspberry Pi, Arduino or other compatible hardware.
  • Data Logging: Detection events and measurements can be stored in CSV and log files for later analysis.

How the Driver Drowsiness Detection System Works

The system follows a multi-stage pipeline. The camera first captures the driver's face, after which MediaPipe detects facial landmarks. The system then extracts measurements related to the eyes, mouth and head position.

These measurements are analysed over time rather than relying on a single video frame. This is important because a normal blink should not be treated as drowsiness, while prolonged eye closure may indicate a micro-sleep event.

The overall processing pipeline is:

  1. Camera
  2. Pre-processing
  3. Face Detection
  4. Facial Landmarks
  5. EAR/MAR/Head Pose
  6. Temporal Analysis
  7. Drowsiness Score
  8. Decision Logic
  9. Alarm

Key Technologies Used

  • Python – Core programming language.
  • OpenCV – Camera handling, image processing and dashboard rendering.
  • MediaPipe – Real-time facial landmark detection.
  • NumPy – Numerical calculations and geometric measurements.
  • PyYAML – Configuration management.
  • Pygame – Optional computer audio alarm.
  • PySerial – Communication between the computer and Arduino.
  • Raspberry Pi GPIO – Optional physical buzzer or relay control.

Step-by-Step Guide to Setting Up the Project

1. Get the Project

Clone or download the project repository and open a terminal inside the project directory.

Terminal
cd drowsiness_detection

The project is organised into separate modules so that individual components such as the camera, landmark detector, scoring system and alarm hardware can be modified independently.

2. Create a Python Virtual Environment

Creating a virtual environment keeps the project's dependencies isolated from other Python applications.

Terminal
python -m venv venv

On Windows:

Windows
venv\Scripts\activate

On macOS, Linux or Raspberry Pi:

macOS, Linux or Raspberry Pi
source venv/bin/activate

3. Install the Required Dependencies

Install the dependencies specified in the project's requirements file:

Terminal
pip install --upgrade pip
pip install -r requirements.txt

The main dependencies include MediaPipe, OpenCV, NumPy and PyYAML. Pygame and PySerial are optional depending on whether a computer speaker or Arduino alarm is being used.

4. Test the Installation

Before connecting a camera, the project's automated tests can be executed to verify the detection and mathematical logic.

Terminal
python -m unittest discover -s tests -t . -v

The project also includes a camera-free simulation that runs the complete decision-making logic using a scripted driving scenario:

Terminal
python tools/simulate_drive.py

5. Connect and Test the Camera

A standard USB webcam or compatible laptop camera can be used. The camera should be positioned approximately 40–80 cm from the driver, preferably around eye level.

The project includes a camera testing utility that can identify available camera indexes and display information such as resolution, FPS and brightness.

Terminal
python tools/test_camera.py

For repeatable demonstrations, a recorded video can also be used instead of a live camera.

Terminal
python main.py --source my_test.mp4

6. Calibrate the Driver

When the system starts, it can perform an automatic five-second calibration period. During calibration, the driver should sit normally, look towards the camera and keep their eyes open with their mouth closed.

The system uses this period to estimate the driver's normal:

  • Open-eye EAR
  • Resting mouth position
  • Neutral head pitch
  • Neutral head yaw

This allows the system to adapt its thresholds to individual drivers rather than relying entirely on fixed values.

The driver can also press c to recalibrate when necessary, such as after changing glasses or moving the camera.

Understanding Eye Aspect Ratio

One of the most important measurements in the system is the Eye Aspect Ratio (EAR). EAR represents the relationship between the vertical opening and horizontal width of the eye.

Formula
EAR = (|p2 − p6| + |p3 − p5|) / (2 × |p1 − p4|)

When the eye is open, the vertical distance between the eyelids is relatively large, producing a higher EAR. When the eye closes, this distance becomes smaller and the EAR decreases.

For example, an eye that is 30 pixels wide with an average vertical opening of 9 pixels produces:

Worked example: open eye
EAR = (9 + 9) / (2 × 30)
    = 0.30

A closed eye with a vertical opening of approximately 3 pixels produces:

Worked example: closed eye
EAR = (3 + 3) / 60
    = 0.10

The project averages measurements from both eyes and uses calibration to determine a personalised eye-closure threshold.

Detecting Yawning with MAR

Eye closure alone is not sufficient to determine drowsiness. The system therefore also monitors the driver's mouth using the Mouth Aspect Ratio (MAR).

Formula
MAR = (|p2 − p8| + |p3 − p7| + |p4 − p6|)
      / (2 × |p1 − p5|)

A closed mouth generally produces a low MAR value, while a wide-open mouth produces a significantly higher value.

However, simply detecting an open mouth is not enough because the driver could be talking. To reduce false detections, the system requires the mouth to remain above the configured MAR threshold for a minimum period before classifying the event as a yawn.

By default, a yawn requires approximately 1.5 seconds of sustained mouth opening.

Head Pose Detection

The system also estimates the driver's head orientation using facial landmarks and OpenCV's solvePnP algorithm.

The estimated head pose is divided into three rotational components:

  • Pitch – movement of the head up and down.
  • Yaw – movement of the head from side to side.
  • Roll – tilting the head sideways.

The system compares these measurements with the driver's calibrated neutral position. A sustained downward head movement can indicate that the driver is becoming drowsy.

Similarly, when the driver turns their head significantly to the side, the system temporarily treats eye measurements as unreliable because the eyes may no longer be visible from a suitable angle.

Temporal Analysis

A major feature of the project is that it does not make decisions from a single frame. Instead, measurements are analysed over time.

The system tracks events such as:

  • Normal blinks
  • Long eye closures
  • Micro-sleeps
  • PERCLOS
  • Reduced blinking
  • Yawning
  • Head drops and nods
  • Face detection loss
  • Looking away

This temporal approach helps distinguish normal behaviour from prolonged events that may indicate fatigue.

Calculating the Drowsiness Score

The system combines multiple indicators into a single drowsiness score.

Scoring model
ACUTE = Eye closure score
       + Head-drop score
       + Yawn-in-progress score
CUMULATIVE = PERCLOS score
            + Yawn history
            + Nod history
            + Long blinks
            + Low blink rate
TOTAL = ACUTE + CUMULATIVE

The acute component represents events happening at the current moment, while the cumulative component represents fatigue-related behaviour observed over a longer period.

The system uses a configurable threshold, with the default alarm threshold set to 100.

This means that a single normal blink cannot normally activate the alarm. Instead, prolonged eye closure or a combination of several fatigue indicators can increase the score until the alarm threshold is reached.

Alarm and Hardware Integration

The project uses an alarm abstraction layer, allowing different hardware outputs to be connected without changing the core drowsiness-detection algorithm.

Available alarm backends include:

  • Simulated Alarm – prints the alarm state to the terminal.
  • Audio Alarm – uses a laptop or external speaker.
  • Raspberry Pi GPIO – controls a buzzer or relay.
  • Arduino – communicates with an Arduino through USB serial.

For example, the application can be started with multiple alarm outputs:

Terminal
python main.py --alarm simulated,audio

For a Raspberry Pi setup:

Raspberry Pi
python main.py --alarm simulated,gpio

The Arduino implementation currently uses a simple serial protocol where A activates the alarm and S deactivates it.

Running the Complete System

Once the camera and required alarm hardware are configured, the complete application can be started with:

Terminal
python main.py

The dashboard displays the live camera feed along with facial landmarks, EAR, MAR, head orientation, drowsiness score, system state and FPS.

The available states include:

  • ALERT
  • BLINKING
  • EYES CLOSED
  • YAWNING
  • HEAD DROP
  • LOOKING AWAY
  • NO FACE
  • DROWSY

During an alarm event, the system displays a warning such as:

On-screen warning
DROWSINESS DETECTED - TAKE A BREAK

Testing the Detection System

The project includes a structured test procedure for evaluating individual behaviours.

TestActionExpected Result
Normal drivingBlink naturally for 30 secondsNo alarm
Quick blinkingBlink repeatedlyBLINKING state without alarm
Short eye closureClose eyes for approximately 1 secondScore increases without immediate alarm
Prolonged closureClose eyes for approximately 3 secondsDrowsiness alarm
YawningKeep mouth open for more than the configured durationYAWNING state
Looking awayTurn head to the sideLOOKING AWAY
Face lossCover the cameraNO FACE / DRIVER NOT VISIBLE
Head dropLower the head for several secondsHEAD DROP followed by alarm

The project also records measurements in CSV format, making it possible to plot EAR, MAR and drowsiness score against time during experiments.

Reducing False Alarms

A driver monitoring system needs to distinguish between normal behaviour and fatigue-related behaviour. Several mechanisms have therefore been incorporated into the detection logic.

  • Duration-based detection: A single closed-eye frame cannot trigger an alarm.
  • Blink filtering: Normal short blinks contribute little or nothing to the eye-closure score.
  • Temporal history: Previous yawns and blinks contribute only within controlled limits.
  • Face-loss handling: Losing the driver's face does not automatically count as drowsiness.
  • Head orientation filtering: Eye measurements are ignored when the driver turns too far away.
  • Alarm hysteresis: Separate alarm activation and deactivation conditions prevent rapid alarm switching.
  • Minimum alarm duration: The alarm remains active for a minimum configured period.

Project Architecture

The project follows a modular architecture so that individual components can be replaced or upgraded independently.

System architecture
Camera
   ↓
Pre-processing
   ↓
MediaPipe Face Detection
   ↓
Facial Landmarks
   ↓
EAR + MAR + Head Pose
   ↓
Temporal Analysis
   ↓
Drowsiness Scoring
   ↓
Decision Logic
   ↓
Alarm Manager
   ├── Speaker
   ├── Raspberry Pi GPIO
   ├── Arduino
   └── Simulated Alarm
                ↓
           Dashboard
                ↓
          Session Logging

The modular design also makes it possible to replace the rule-based detection system with a machine-learning model in the future without rebuilding the entire application.

Logging and Data Analysis

Every monitoring session can generate timestamped log and CSV files. The CSV contains information such as:

CSV columns
timestamp
ear
mar
pitch
yaw
perclos
blink_rate
score
acute
cumulative
state
alarm

This data can be imported into Excel, Python or other analysis tools to visualise how eye closure, yawning and head movement affect the drowsiness score over time.

Raspberry Pi Deployment

The system is designed to run not only on a desktop computer but also on a Raspberry Pi 4 or Raspberry Pi 5.

A Raspberry Pi deployment can use a Pi camera, USB camera and GPIO-connected buzzer or relay. For better real-time performance, the system can operate at resolutions such as 640×480 or 480×360.

A typical Raspberry Pi command is:

Raspberry Pi
python main.py --alarm simulated,gpio

Troubleshooting Common Issues

  • Camera not detected: Run python tools/test_camera.py and check the available camera indexes.
  • Low FPS: Reduce the camera resolution, disable unnecessary preprocessing and close other applications.
  • MediaPipe compatibility error: Use a supported MediaPipe version or download the model required by the Tasks API.
  • False eye-closure detection: Recalibrate the system while looking directly at the camera.
  • Yawning detected while talking: Increase the minimum yawn duration or adjust the MAR threshold.
  • Head-drop detection reversed: Check the pitch direction and enable the appropriate pitch inversion setting.
  • Frequent NO FACE warnings: Improve lighting and ensure the entire face remains visible.
  • No alarm sound: Test the selected alarm backend using tools/test_alarm.py.

Limitations of the Current System

This project is a rule-based driver-monitoring prototype intended for education, experimentation and demonstration. The automated tests verify that the implemented logic behaves according to its design, but they do not establish real-world drowsiness-detection accuracy across different drivers and environments.

Several limitations remain:

  • Multiple faces are currently handled by selecting the largest detected face.
  • Dark sunglasses can make eye measurements unreliable.
  • Extreme lighting and camera angles can affect facial landmark detection.
  • Thresholds and scoring weights require evaluation across different subjects.
  • The Arduino protocol currently does not provide hardware acknowledgements.
  • Successful transmission of an alarm command does not necessarily confirm physical alarm activation.
  • Camera failure and recovery require careful handling to prevent stale monitoring states.

For these reasons, the system should be treated as a driver-assistance prototype rather than a certified automotive safety device.

Future Improvements: CNN and LSTM

The current geometric approach is lightweight and relatively easy to interpret, but fixed thresholds cannot capture every variation in human faces and driving conditions.

A future version could incorporate deep learning models to improve adaptability.

1. CNN-Based Eye State Detection

A lightweight CNN such as MobileNetV3-Small could classify eye crops as open or closed. This model could replace or complement EAR-based eye detection.

2. LSTM or GRU-Based Temporal Analysis

An LSTM, GRU or Temporal CNN could process sequences containing EAR, MAR, pitch, yaw, blink duration and other features over several seconds. Instead of relying entirely on manually selected thresholds, the model could learn temporal patterns associated with increasing fatigue.

3. CNN-LSTM Architecture

A more advanced implementation could combine facial image features from a CNN with temporal modelling through an LSTM or GRU. This would allow both facial appearance and changes over time to be considered.

4. Edge Deployment

The final model could be converted to TensorFlow Lite or ONNX and quantised for deployment on Raspberry Pi or other edge-computing hardware.

5. Additional Vehicle Signals

Future versions could combine camera-based monitoring with vehicle information such as steering behaviour, lane-keeping data and CAN bus signals. Combining independent indicators could provide a more comprehensive driver-monitoring system.

Conclusion

The Driver Drowsiness Detection and Alarm System demonstrates how computer vision can be combined with temporal analysis and hardware interfaces to create a real-time driver-monitoring prototype.

Instead of relying on a single signal, the system analyses several indicators including eye closure, blinking, yawning and head movement. These measurements are converted into a configurable drowsiness score and used to activate an alarm when prolonged or combined fatigue-related behaviour is detected.

The modular architecture also provides a foundation for future development, including CNN-based eye-state detection, LSTM-based temporal modelling, infrared cameras, vehicle CAN integration and edge-AI deployment.

Project

Author: Dhruvraj Singh Shekhawat

Technologies: Python, OpenCV, MediaPipe, NumPy, Raspberry Pi, Arduino, Computer Vision

Comments

1000 characters left. Plain text; links are allowed but limited.

Loading comments…

    Suggest a blog topic