1 ST-TS-HGR-NET
ST-TS-HGR-Net: Spatial-Temporal and Temporal-Spatial Hand Gesture Recognition Network
ST-TS-HGR-NET is a two-branch hand-gesture network that combines spatial-then-temporal and temporal-then-spatial Gaussian aggregation before learning a final SPD representation.
Ref: @nguyenNeuralNetworkBased2019
Code not available.
Overview
ST-TS-HGR-NET directly accepts sequences of 3-D hand-joint coordinates and contains three components:
CONV: shared per-frame spatial convolution over a hand-skeleton grid.ST-GA-NETandTS-GA-NET: complementary Gaussian aggregation branches that encode first- and second-order statistics.SPDC-NET: learned SPD aggregation, logarithmic mapping, vectorization, and classification.
Unlike SPDNet, which receives a covariance matrix computed from a whole sequence, ST-TS-HGR-NET learns from raw joint coordinates, keeps mean as well as covariance information, and models separate fingers and temporal partitions at finer granularity.




Architecture
Spatial Convolution
The irregular hand skeleton is mapped to a 2-D grid with three channels for , , and . Some joints are removed and additional cross-finger neighborhood connections are introduced. Each node has at most nine neighbors, including itself. For joint at frame ,
where , is the grid neighborhood, and one of nine shared weight matrices is selected by relative grid position. The same filters are shared across all frames.
Gaussian Embedding (GaussAgg)
To understand the SPDAgg layers, let us follow the authors’ explanation:
- Let and be respectively the number of hand joints and the length of the skeleton sequence. Let us denote by , the 3D coordinates of hand joint at frame .
- Let be the output dimension of the convolutional layer.
- Let us denote by , the output of the convolutional layer. The output feature vector at node i is computed as:
where is the set of neighbors of node , is the filter weight matrix.
- To aggregate features in a branch associated with sub-sequence and finger , , each frame of sub-sequence is processed through 4 layers.
- Let be the set of hand joints belonging to finger , be the beginning and ending frames of sub-sequence , a given frame of sub-sequence , the subset of output feature vectors of the convolutional layer fed to the branch.
- Consider a sliding window centered on frame . Following previous works, are assumed i.i.d. samples from a Gaussian distribution:
with parameters estimated as
- Based on the method in @lovricMultivariateNormalDistributions2000 that embeds the space of Gaussians in the Riemannian symmetric space, the Gaussian can be identified as a SPD matrix given by:
- The GaussAgg layer performs this computation:
This representation preserves both first-order and second-order statistics.
Branch Design
Each sequence yields six temporal selections: the complete sequence, two halves, and three thirds. Pairing these with five fingers gives 30 branches in ST-GA-NET and 30 branches in TS-GA-NET.
- ST branch: for each subsequence/finger/frame, the first GaussAgg aggregates every joint of that finger over the sliding window; the SPD matrix passes through ReEig → LogEig →
VecMat(which preserves the Frobenius inner product by multiplying off-diagonal upper-triangular entries by ). A second GaussAgg then aggregates these vectors over all frames of the subsequence, producing an SPD matrix that describes the temporal variation of finger within subsequence . - TS branch: each subsequence is divided into equal temporal pieces. For every joint and piece, the first GaussAgg computes an SPD matrix from that joint’s temporal samples; after ReEig → LogEig → VecMat, the second GaussAgg aggregates the vectors jointly over temporal pieces and the joints of finger . This reverses the aggregation order: temporal statistics first per joint, then spatial pooling over the finger.
SPD Aggregation and Classification
The 60 branch outputs are fused by SPDAgg:
with . Positive definiteness follows when has full row rank; it is constrained to the compact Stiefel manifold and updated by projected gradient descent followed by retraction. The tail is SPDAgg → LogEig → upper-triangular vectorization (off-diagonals ×), and the reported final classifier is a LIBLINEAR SVM.
Model Parameters
- Spatial-convolution output dimension: ; maximum grid neighborhood: 9 nodes.
- Temporal selections per sequence: 6; fingers: 5; branches: 30 + 30 = 60.
- Sliding-window radius (3 consecutive frames); TS subdivisions ; sequences normalized to frames.
- First GaussAgg: 9-D vectors → SPD; VecMat → 55 features; second GaussAgg → SPD.
- SPDAgg: input , output ; each transformation matrix .
- ReEig threshold: .
- DHG input: 22 hand joints; FPHA input: 21 hand joints (no palm).
Training Parameters
- Batch size: 30.
- Learning rate: 0.01.
- Selected network checkpoint: epoch 15.
- Optimizer name, momentum, weight decay: not reported.
- SVM: LIBLINEAR L2-regularized L2-loss dual; ; stopping tolerance 0.1; no bias term.
- Implementation: MATLAB R2015b CPU; FPHA ≈22 min per training epoch and 7 min per test epoch.
- Frame counts 300/500/800 tested; differences marginal.
Results
Main results (accuracy %):
| Dataset/protocol | ST-TS-HGR-NET | SPDNet |
|---|---|---|
| DHG, fixed split, 14 gestures | 94.29 | 75.24 |
| DHG, fixed split, 28 gestures | 89.40 | 69.64 |
| DHG, LOSO, 14 gestures | 87.3 | - |
| DHG, LOSO, 28 gestures | 83.4 | - |
| FPHA | 93.22 | 84.35 |
Gains over SPDNet: +19.05 / +19.76 pp (fixed DHG 14/28) and +8.87 pp (FPHA). FPHA baselines: Gram Matrix 85.39, Lie Group 82.69, LSTM 80.14, ST-GCN 81.30, Grassmann network 77.57.
Ablations (FPHA / DHG-14 / DHG-28):
| Ablation | FPHA | DHG-14 | DHG-28 |
|---|---|---|---|
| ST-HGR-NET only | 91.83 | 93.21 | 89.29 |
| TS-HGR-NET only | 90.96 | 93.33 | 88.21 |
| Full ST-TS-HGR-NET | 93.22 | 94.29 | 89.40 |
| 3 neighbors (physical) | 91.65 | 93.10 | 88.33 |
| 9 neighbors (augmented grid) | 93.22 | 94.29 | 89.40 |
| 93.22 | 94.29 | 89.40 | |
| 93.04 | 94.17 | 89.04 | |
| 93.04 | 94.29 | 89.40 | |
| 93.33 | 94.29 | 89.40 | |
| 92.87 | 94.05 | 88.93 | |
| 92.70 | 94.29 | 89.04 |