HamburgerMenu
hirist

SLAM & Computer Vision Engineer - YOLO

Technosoft Engineering Projects Limited
5 - 10 Years
Multiple Locations

Posted on: 17/08/2026

Job Description

ABOUT US :

Technosoft Engineering Solutions is building an AI-powered visual intelligence platform for the Indian construction industry. We bridge the gap between what is planned (BIM models, 2D drawings) and what is built on site, using real-time visual intelligence and spatial computing.

THE ROLE :

We are looking for a Computer Vision & AI Engineer to design and build scalable vision AI pipelines that process large volumes & variety. You will implement end-to-end AI pipeline : video preprocessing, model architecture (YOLO, SAM, segmentation models), training on construction datasets, inference optimization, and MLOps for continuous improvement.

WHAT WE ARE LOOKING FOR :

Computer Vision & Deep Learning (Core Requirement) :

- Hands-on experience building object detection and segmentation models in production : YOLO, SAM (Segment Anything Model), or similar architectures

- Strong fundamentals in CNNs, vision transformers, attention mechanisms, and multi-scale feature extraction

- Practical experience with semantic segmentation and instance segmentation for complex, cluttered environments

- Model training on custom datasets : data annotation pipelines, class imbalance handling, augmentation strategies

- Transfer learning and fine-tuning : adapting pre-trained models to domain-specific tasks

- Experience with 2D - to - 3D or video-to-BIM alignment is a strong plus

Large-Scale Video Processing & Inference Optimization (Core Requirement) :

- Experience building scalable video processing pipelines handling TBs of data per month

- Batch processing architecture : distributed inference across GPU clusters, queueing systems (RabbitMQ, Kafka, etc.), chunked video processing

- Model optimization for production : quantization, pruning, ONNX Runtime, TensorRT

- Frame sampling strategies for efficient video analysis (every Nth frame, keyframe extraction, motion-based sampling)

- Experience with cloud-based GPU inference (AWS EC2 P3/P4, Lambda, SageMaker) and cost optimization

Training Data & Domain-Specific Model Development (Core Requirement) :

- Experience training models on noisy, real-world data : handling dust, shadows, occlusions, varying lighting conditions, motion blur

- Building custom datasets : defining labeling guidelines, managing annotation teams, quality control

- Multi-class classification with 50+ element classes across various construction segments

- Defect detection modelling - cracks, rebar exposure, missing components, misalignments, incomplete work is strong plus

Hybrid 2D/3D AI & BIM Integration (Good to have) :

- Experience working with BIM models (IFC files) or CAD drawings as reference for AI analysis

- Spatial reasoning : matching detected elements in video frames to BIM element locations using coordinate transformations

- Zone-based analysis : segmenting floor plans into zones, aggregating detection results per zone/floor

- Progress quantification : computing completion percentages from detection outputs (e.g., "Floor 7 Finishes 40% complete")

MLOps & Production AI Systems (Good to have) :

- Model versioning, experiment tracking (MLflow, Weights & Biases, ClearML)

- CI/CD for ML : automated retraining pipelines, A/B testing new models in production, monitoring model drift

- Production monitoring : precision/recall tracking per class, confidence score distributions, detection latency, inference cost per video

- Human-in-the-loop (HITL) workflows : flagging low-confidence predictions for manual review, feedback loops for continuous improvement

- Experience with model explainability (Grad-CAM, SHAP) for debugging and trust

Programming Languages :

- Python - expert level (PyTorch, TensorFlow, OpenCV, NumPy, Pandas)

- Deep learning frameworks : PyTorch (preferred) or TensorFlow/Keras

- Computer vision libraries : OpenCV, Albumentations, Pillow

- GPU programming : CUDA basics, understanding GPU memory management

- Familiar with video codecs (H.264, H.265), FFmpeg for video manipulation

QUALIFICATIONS :

- 5 - 10 years of professional software engineering experience

- 3 - 5 years building production computer vision / deep learning system

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...