| 項次 | 料號 | 數量 | 單位 | 需用日 | 小計 |
| | 品名規格 | | | 單價 | |
| 規格 | |
| 10 |
|
1 |
CS |
20261130 |
|
| | 遮蔽感知多模態模型訓練技術 | | |
________________ |
________________ |
| | 【Item I】 Development of an Occlusion-Aware Multimodal Intelligent Inference Model
Project Background: Reliable multimodal perception and inference are critical for applicatio ns in which fine-grained hand motion must remain stable under self-occlu sion, two-hand interaction, and complex foreground/background relationsh ips. The planned R&<)>D work focuses on estimating continuous surface -visibility and confidence values for 3D hand keypoints by combining ima ge-level information, 3D hand geometry, camera projection information, d epth-buffer reasoning, and features extracted from a WiLoR-based hand re construction pipeline. The resulting multimodal representation is intend ed to distinguish keypoints that are fully visible, self-occluded, occlu ded by the other hand, outside the image, or otherwise invalid, and to p rovide a richer hand-state condition than a conventional single-source 2 D coordinate representation.
Project Objectives: Construct an InterHand2.6M-Based Multimodal Visibility Dataset: Use camera parameters, image annotations, 3D keypoint ground truth, and fitted MANO parameters to reconstruct the geometric relationship among hand keypoints, hand surfaces, image observations, and the camera view. Generate visibility-related training targets using projected 3D geometr y and depth/surface reasoning.
Develop a Multimodal Visibility Inference Head: Freeze the WiLoR backbone as the primary feature extractor and train a l ightweight inference head to estimate continuous visibility values for h and keypoints from the learned visual and geometric representation.
Evaluate Regression and Ranking Objectives: Use regression as the primary learning objective and experimentally comp are it with a regression-plus-ranking formulation. The comparison shall determine whether explicit keypoint-order supervision improves or degra des visibility inference under the available data distribution.
Validate Occlusion-Aware Inference Robustness: Evaluate model behavior across single-hand and interacting-hand cases, i ncluding different overlap strengths, with particular attention to self- occlusion and cross-hand occlusion.
Establish Reproducible Evaluation Results: Report error, correlation, rank-consistency, and pairwise-order metrics, together with representative visual examples and failure cases, to supp ort subsequent integration into multimodal hand-state-conditioned genera tion workflows.
Project Scope & Experimental Steps 1. Multimodal Dataset Preparation and Visibility Target Construction Objective Prepare training and validation data that jointly describe image observa tions, 3D hand geometry, camera relationships, and occlusion-aware visib ility targets. Tasks: Load InterHand2.6M camera parameters and convert world-space 3D coordina tes into camera and image coordinates. Use image-level metadata to preserve capture, camera, frame, and left/ri ght-hand state information. Use 3D keypoint annotations as the geometric ground truth. Use fitted MANO parameters to reconstruct complete 3D hand meshes for su rface-aware visibility analysis. Apply depth-buffer / Z-buffer reasoning to determine foreground/backgrou nd relationships and surface visibility. Represent keypoint status using the available categories: visible, self- occluded, occluded by the other hand, outside the image, and invalid.
2. Multimodal Feature Extraction and Intelligent Inference Model Objective Build an occlusion-aware visibility inference model on top of the existi ng WiLoR hand reconstruction pipeline while minimizing changes to the ba ckbone and retaining its learned visual-geometric representation. Tasks: Use WiLoR-derived visual and geometric hand features as the model input representation. Freeze the WiLoR backbone parameters during the multimodal visibility-in ference experiments. Attach a lightweight inference head that outputs continuous per-keypoint visibility scores. Preserve the WiLoR geometry path, including coarse MANO prediction and m ulti-scale refinement, as the underlying hand-reconstruction feature sou rce. Save model checkpoints and the corresponding experiment configuration fo r each major training variant.
3. Inference Objective Design and Training Experiments Objective Determine whether continuous regression alone or regression combined wit h a ranking objective provides the most reliable visibility inference. Tasks: Train a regression-only baseline using MAE-oriented continuous visibilit y supervision. Train a regression + ranking variant so that the model also learns relat ive visibility ordering among keypoints. Compare the ranking-aware model against a longer-training regression-onl y control under the same evaluation protocol. Treat regression-only training as the primary reference track unless ran king demonstrates a consistent improvement on full validation and intera cting-hand subsets. Analyze whether the ranking term introduces noise when the training dist ribution is dominated by single-hand images.
4. Intelligent Inference Evaluation and Performance Testing Objective Quantitatively assess inference accuracy, correlation with ground truth, and consistency of relative visibility ordering. Tasks: Regression Error: Mean Absolute Error (MAE) and, where applicable, Root Mean Square Error (RMSE). Linear Correlation: Pearson correlation. Rank Correlation: macro Spearman correlation and macro Kendall correlati on. Relative Ordering: pair accuracy for keypoint-visibility comparisons. Use global-mean and per-keypoint-mean predictors as simple reference bas elines.
5. Cross-Hand Occlusion and Multimodal Robustness Analysis Objective Verify whether the multimodal visibility inference remains stable when t he two hands overlap and one hand occludes the other. Tasks: Group interacting-hand samples by overlap strength: no overlap, weak, mo derate, and strong. Compare regression-only and regression + ranking variants using cross-ha nd pair accuracy. Inspect cases in which keypoints are occluded by the other hand and comp are them with self-occlusion cases. Document whether performance degradation is associated with overlap stre ngth, visibility ambiguity, projection error, or incomplete geometric la bels. Include representative high-confidence and lower-confidence prediction e xamples for qualitative review.
6. Analysis, Integration Guidance, and Deliverables Objective Deliver a reproducible technical package that explains multimodal data c onstruction, inference-model design, experimental findings, and the reco mmended model variant. Tasks: Training Specification: frozen-backbone configuration, inference-head de sign, loss functions, split definitions, and checkpoint settings. Evaluation Report: complete metric tables for baseline, regression-only, and regression + ranking variants. Occlusion Analysis: results for self-occlusion, cross-hand occlusion, an d overlap-strength groups. Qualitative Results: representative hand-state overlays and predicted vi sibility examples, including successful and difficult cases. Model Assets: trained multimodal visibility-inference head/checkpoint an d the scripts or configuration required to reproduce inference and evalu ation.
|
|