Building a Driver Drowsiness Detection and Alarm System
A real-time, camera-based prototype that tracks eye closure, yawning and head movement with Python, OpenCV and MediaPipe, and sounds an alarm when fatigue builds up.
In this post
Driver fatigue is one of the major challenges in road safety, particularly during long journeys and low-attention driving conditions. A driver may experience prolonged eye closure, frequent yawning, reduced blinking, or involuntary head movements before reaching a dangerous level of drowsiness. Detecting these signs early can provide an opportunity to alert the driver before the situation becomes critical.
The Driver Drowsiness Detection and Alarm System is a real-time, camera-based driver monitoring prototype that uses computer vision and facial landmarks to identify potential signs of driver fatigue. The system uses MediaPipe for facial landmark detection, calculates Eye Aspect Ratio (EAR) and Mouth Aspect Ratio (MAR), estimates head pose, and analyses these measurements over time to determine a drowsiness score.
When prolonged drowsiness is detected, the system activates an alarm through a computer speaker, Raspberry Pi GPIO, Arduino, or a simulated software alarm. The project can run on a regular laptop as well as on a Raspberry Pi 4/5.
Why Use a Driver Drowsiness Detection System?
Traditional driver monitoring often relies on the driver's own awareness of fatigue. However, signs such as prolonged eye closure, yawning and head dropping can occur before the driver consciously recognises that they are becoming drowsy.
A camera-based monitoring system provides a non-contact method of continuously observing these indicators. Some of the main advantages of this approach include:
- Real-Time Monitoring: The driver's face can be analysed continuously using a standard camera.
- Non-Invasive Detection: The system does not require wearable sensors or physical contact with the driver.
- Multiple Indicators: Eye closure, blinking, yawning and head movement are analysed together.
- Personalised Calibration: The system learns the driver's normal eye and head position during startup.
- Hardware Flexibility: The alarm can operate through a laptop, Raspberry Pi, Arduino or other compatible hardware.
- Data Logging: Detection events and measurements can be stored in CSV and log files for later analysis.
How the Driver Drowsiness Detection System Works
The system follows a multi-stage pipeline. The camera first captures the driver's face, after which MediaPipe detects facial landmarks. The system then extracts measurements related to the eyes, mouth and head position.
These measurements are analysed over time rather than relying on a single video frame. This is important because a normal blink should not be treated as drowsiness, while prolonged eye closure may indicate a micro-sleep event.
The overall processing pipeline is:
- Camera
- Pre-processing
- Face Detection
- Facial Landmarks
- EAR/MAR/Head Pose
- Temporal Analysis
- Drowsiness Score
- Decision Logic
- Alarm
Key Technologies Used
- Python – Core programming language.
- OpenCV – Camera handling, image processing and dashboard rendering.
- MediaPipe – Real-time facial landmark detection.
- NumPy – Numerical calculations and geometric measurements.
- PyYAML – Configuration management.
- Pygame – Optional computer audio alarm.
- PySerial – Communication between the computer and Arduino.
- Raspberry Pi GPIO – Optional physical buzzer or relay control.
Step-by-Step Guide to Setting Up the Project
1. Get the Project
Clone or download the project repository and open a terminal inside the project directory.
cd drowsiness_detectionThe project is organised into separate modules so that individual components such as the camera, landmark detector, scoring system and alarm hardware can be modified independently.
2. Create a Python Virtual Environment
Creating a virtual environment keeps the project's dependencies isolated from other Python applications.
python -m venv venvOn Windows:
venv\Scripts\activateOn macOS, Linux or Raspberry Pi:
source venv/bin/activate3. Install the Required Dependencies
Install the dependencies specified in the project's requirements file:
pip install --upgrade pip
pip install -r requirements.txtThe main dependencies include MediaPipe, OpenCV, NumPy and PyYAML. Pygame and PySerial are optional depending on whether a computer speaker or Arduino alarm is being used.
4. Test the Installation
Before connecting a camera, the project's automated tests can be executed to verify the detection and mathematical logic.
python -m unittest discover -s tests -t . -vThe project also includes a camera-free simulation that runs the complete decision-making logic using a scripted driving scenario:
python tools/simulate_drive.py5. Connect and Test the Camera
A standard USB webcam or compatible laptop camera can be used. The camera should be positioned approximately 40–80 cm from the driver, preferably around eye level.
The project includes a camera testing utility that can identify available camera indexes and display information such as resolution, FPS and brightness.
python tools/test_camera.pyFor repeatable demonstrations, a recorded video can also be used instead of a live camera.
python main.py --source my_test.mp46. Calibrate the Driver
When the system starts, it can perform an automatic five-second calibration period. During calibration, the driver should sit normally, look towards the camera and keep their eyes open with their mouth closed.
The system uses this period to estimate the driver's normal:
- Open-eye EAR
- Resting mouth position
- Neutral head pitch
- Neutral head yaw
This allows the system to adapt its thresholds to individual drivers rather than relying entirely on fixed values.
The driver can also press c to recalibrate when necessary, such as after changing glasses or moving the camera.
Understanding Eye Aspect Ratio
One of the most important measurements in the system is the Eye Aspect Ratio (EAR). EAR represents the relationship between the vertical opening and horizontal width of the eye.
EAR = (|p2 − p6| + |p3 − p5|) / (2 × |p1 − p4|)When the eye is open, the vertical distance between the eyelids is relatively large, producing a higher EAR. When the eye closes, this distance becomes smaller and the EAR decreases.
For example, an eye that is 30 pixels wide with an average vertical opening of 9 pixels produces:
EAR = (9 + 9) / (2 × 30)
= 0.30A closed eye with a vertical opening of approximately 3 pixels produces:
EAR = (3 + 3) / 60
= 0.10The project averages measurements from both eyes and uses calibration to determine a personalised eye-closure threshold.
Detecting Yawning with MAR
Eye closure alone is not sufficient to determine drowsiness. The system therefore also monitors the driver's mouth using the Mouth Aspect Ratio (MAR).
MAR = (|p2 − p8| + |p3 − p7| + |p4 − p6|)
/ (2 × |p1 − p5|)A closed mouth generally produces a low MAR value, while a wide-open mouth produces a significantly higher value.
However, simply detecting an open mouth is not enough because the driver could be talking. To reduce false detections, the system requires the mouth to remain above the configured MAR threshold for a minimum period before classifying the event as a yawn.
By default, a yawn requires approximately 1.5 seconds of sustained mouth opening.
Head Pose Detection
The system also estimates the driver's head orientation using facial landmarks and OpenCV's solvePnP algorithm.
The estimated head pose is divided into three rotational components:
- Pitch – movement of the head up and down.
- Yaw – movement of the head from side to side.
- Roll – tilting the head sideways.
The system compares these measurements with the driver's calibrated neutral position. A sustained downward head movement can indicate that the driver is becoming drowsy.
Similarly, when the driver turns their head significantly to the side, the system temporarily treats eye measurements as unreliable because the eyes may no longer be visible from a suitable angle.
Temporal Analysis
A major feature of the project is that it does not make decisions from a single frame. Instead, measurements are analysed over time.
The system tracks events such as:
- Normal blinks
- Long eye closures
- Micro-sleeps
- PERCLOS
- Reduced blinking
- Yawning
- Head drops and nods
- Face detection loss
- Looking away
This temporal approach helps distinguish normal behaviour from prolonged events that may indicate fatigue.
Calculating the Drowsiness Score
The system combines multiple indicators into a single drowsiness score.
ACUTE = Eye closure score
+ Head-drop score
+ Yawn-in-progress score
CUMULATIVE = PERCLOS score
+ Yawn history
+ Nod history
+ Long blinks
+ Low blink rate
TOTAL = ACUTE + CUMULATIVEThe acute component represents events happening at the current moment, while the cumulative component represents fatigue-related behaviour observed over a longer period.
The system uses a configurable threshold, with the default alarm threshold set to 100.
This means that a single normal blink cannot normally activate the alarm. Instead, prolonged eye closure or a combination of several fatigue indicators can increase the score until the alarm threshold is reached.
Alarm and Hardware Integration
The project uses an alarm abstraction layer, allowing different hardware outputs to be connected without changing the core drowsiness-detection algorithm.
Available alarm backends include:
- Simulated Alarm – prints the alarm state to the terminal.
- Audio Alarm – uses a laptop or external speaker.
- Raspberry Pi GPIO – controls a buzzer or relay.
- Arduino – communicates with an Arduino through USB serial.
For example, the application can be started with multiple alarm outputs:
python main.py --alarm simulated,audioFor a Raspberry Pi setup:
python main.py --alarm simulated,gpioThe Arduino implementation currently uses a simple serial protocol where A activates the alarm and S deactivates it.
Running the Complete System
Once the camera and required alarm hardware are configured, the complete application can be started with:
python main.pyThe dashboard displays the live camera feed along with facial landmarks, EAR, MAR, head orientation, drowsiness score, system state and FPS.
The available states include:
- ALERT
- BLINKING
- EYES CLOSED
- YAWNING
- HEAD DROP
- LOOKING AWAY
- NO FACE
- DROWSY
During an alarm event, the system displays a warning such as:
DROWSINESS DETECTED - TAKE A BREAKTesting the Detection System
The project includes a structured test procedure for evaluating individual behaviours.
| Test | Action | Expected Result |
|---|---|---|
| Normal driving | Blink naturally for 30 seconds | No alarm |
| Quick blinking | Blink repeatedly | BLINKING state without alarm |
| Short eye closure | Close eyes for approximately 1 second | Score increases without immediate alarm |
| Prolonged closure | Close eyes for approximately 3 seconds | Drowsiness alarm |
| Yawning | Keep mouth open for more than the configured duration | YAWNING state |
| Looking away | Turn head to the side | LOOKING AWAY |
| Face loss | Cover the camera | NO FACE / DRIVER NOT VISIBLE |
| Head drop | Lower the head for several seconds | HEAD DROP followed by alarm |
The project also records measurements in CSV format, making it possible to plot EAR, MAR and drowsiness score against time during experiments.
Reducing False Alarms
A driver monitoring system needs to distinguish between normal behaviour and fatigue-related behaviour. Several mechanisms have therefore been incorporated into the detection logic.
- Duration-based detection: A single closed-eye frame cannot trigger an alarm.
- Blink filtering: Normal short blinks contribute little or nothing to the eye-closure score.
- Temporal history: Previous yawns and blinks contribute only within controlled limits.
- Face-loss handling: Losing the driver's face does not automatically count as drowsiness.
- Head orientation filtering: Eye measurements are ignored when the driver turns too far away.
- Alarm hysteresis: Separate alarm activation and deactivation conditions prevent rapid alarm switching.
- Minimum alarm duration: The alarm remains active for a minimum configured period.
Project Architecture
The project follows a modular architecture so that individual components can be replaced or upgraded independently.
Camera
↓
Pre-processing
↓
MediaPipe Face Detection
↓
Facial Landmarks
↓
EAR + MAR + Head Pose
↓
Temporal Analysis
↓
Drowsiness Scoring
↓
Decision Logic
↓
Alarm Manager
├── Speaker
├── Raspberry Pi GPIO
├── Arduino
└── Simulated Alarm
↓
Dashboard
↓
Session LoggingThe modular design also makes it possible to replace the rule-based detection system with a machine-learning model in the future without rebuilding the entire application.
Logging and Data Analysis
Every monitoring session can generate timestamped log and CSV files. The CSV contains information such as:
timestamp
ear
mar
pitch
yaw
perclos
blink_rate
score
acute
cumulative
state
alarmThis data can be imported into Excel, Python or other analysis tools to visualise how eye closure, yawning and head movement affect the drowsiness score over time.
Raspberry Pi Deployment
The system is designed to run not only on a desktop computer but also on a Raspberry Pi 4 or Raspberry Pi 5.
A Raspberry Pi deployment can use a Pi camera, USB camera and GPIO-connected buzzer or relay. For better real-time performance, the system can operate at resolutions such as 640×480 or 480×360.
A typical Raspberry Pi command is:
python main.py --alarm simulated,gpioTroubleshooting Common Issues
- Camera not detected: Run
python tools/test_camera.pyand check the available camera indexes. - Low FPS: Reduce the camera resolution, disable unnecessary preprocessing and close other applications.
- MediaPipe compatibility error: Use a supported MediaPipe version or download the model required by the Tasks API.
- False eye-closure detection: Recalibrate the system while looking directly at the camera.
- Yawning detected while talking: Increase the minimum yawn duration or adjust the MAR threshold.
- Head-drop detection reversed: Check the pitch direction and enable the appropriate pitch inversion setting.
- Frequent NO FACE warnings: Improve lighting and ensure the entire face remains visible.
- No alarm sound: Test the selected alarm backend using
tools/test_alarm.py.
Limitations of the Current System
This project is a rule-based driver-monitoring prototype intended for education, experimentation and demonstration. The automated tests verify that the implemented logic behaves according to its design, but they do not establish real-world drowsiness-detection accuracy across different drivers and environments.
Several limitations remain:
- Multiple faces are currently handled by selecting the largest detected face.
- Dark sunglasses can make eye measurements unreliable.
- Extreme lighting and camera angles can affect facial landmark detection.
- Thresholds and scoring weights require evaluation across different subjects.
- The Arduino protocol currently does not provide hardware acknowledgements.
- Successful transmission of an alarm command does not necessarily confirm physical alarm activation.
- Camera failure and recovery require careful handling to prevent stale monitoring states.
For these reasons, the system should be treated as a driver-assistance prototype rather than a certified automotive safety device.
Future Improvements: CNN and LSTM
The current geometric approach is lightweight and relatively easy to interpret, but fixed thresholds cannot capture every variation in human faces and driving conditions.
A future version could incorporate deep learning models to improve adaptability.
1. CNN-Based Eye State Detection
A lightweight CNN such as MobileNetV3-Small could classify eye crops as open or closed. This model could replace or complement EAR-based eye detection.
2. LSTM or GRU-Based Temporal Analysis
An LSTM, GRU or Temporal CNN could process sequences containing EAR, MAR, pitch, yaw, blink duration and other features over several seconds. Instead of relying entirely on manually selected thresholds, the model could learn temporal patterns associated with increasing fatigue.
3. CNN-LSTM Architecture
A more advanced implementation could combine facial image features from a CNN with temporal modelling through an LSTM or GRU. This would allow both facial appearance and changes over time to be considered.
4. Edge Deployment
The final model could be converted to TensorFlow Lite or ONNX and quantised for deployment on Raspberry Pi or other edge-computing hardware.
5. Additional Vehicle Signals
Future versions could combine camera-based monitoring with vehicle information such as steering behaviour, lane-keeping data and CAN bus signals. Combining independent indicators could provide a more comprehensive driver-monitoring system.
Conclusion
The Driver Drowsiness Detection and Alarm System demonstrates how computer vision can be combined with temporal analysis and hardware interfaces to create a real-time driver-monitoring prototype.
Instead of relying on a single signal, the system analyses several indicators including eye closure, blinking, yawning and head movement. These measurements are converted into a configurable drowsiness score and used to activate an alarm when prolonged or combined fatigue-related behaviour is detected.
The modular architecture also provides a foundation for future development, including CNN-based eye-state detection, LSTM-based temporal modelling, infrared cameras, vehicle CAN integration and edge-AI deployment.
Comments
Loading comments…