Abstract
Especially in scenarios with highly packed gatherings such as religious pilgrimages, this study focuses on the critical requirement for improved crowd-analysis techniques. In these settings that ensure safety and situational awareness, the importance of video monitoring and visual analysis has significantly increased. Despite significant advancements in human pose estimation, challenges remain largely unresolved in highly crowded situations where individual recognition and movement tracking become more complex. Moreover, there are no strong benchmarking instruments that cater to these demanding criteria. We propose a novel method that precisely predicts individual postures in crowded environments to overcome these constraints. Our method identifies and analyses many human postures using a Mask R-CNN architecture with a ResNet101 backbone, therefore facilitating automated crowd behavior detection. In this study, we provide a specialized dataset, HAJJ-Crowd, which comprises annotated video sequences used to assess pose estimation methods in high-density real-world environments. In this data set, our approach was achieved with 78.0 mean average precision (mAP) in this massive crowd domain. The data set is available here https://drive.google.com/drive/folders/1-g-de-9YINLCgObC3XvaoPTCbbI9EI-F.