Passive Optical Networks (PONs) constitute the backbone of modern broadband access due to their high capacity and cost efficiency; however, upstream Dynamic Bandwidth Allocation (DBA) remains a critical challenge under heterogeneous, bursty traffic and strict latency constraints. This paper presents Service-Class Hierarchical Channel-aware Multi-Objective Dynamic Bandwidth Allocation (SC-H-CMO-DBA), a hierarchical reinforcement learning-based DBA framework designed for ITU-T-compliant PON systems operating with 125- \(\mu \) s transmission-convergence frames. The proposed architecture employs a two-level learning structure. At the intra-OLT level, a delay-aware Deep Q-Network (DQN) performs fine-grained, tile-level scheduling among Optical Network Terminals (ONTs) by jointly observing queue backlog, head-of-line delay, and traffic class composition (video, voice, and data). At the higher level, a lightweight PPO-based coordination mechanism regulates long-term scheduling bias across ONTs, stabilizing utilization and mitigating congestion under uneven and overload traffic conditions. The proposed scheme is benchmarked against Greedy O-OFDMA scheduler and flat reinforcement learning approaches, namely Double Deep Q Networks (DDQN-DBA) and Proximal Policy Optimization (PPO)-DBA, under identical capacity and timing constraints. Results show that SC-H-CMO-DBA maintains near-optimal utilization and goodput while achieving lower mean delay, P95 tail delay, jitter, and packet loss than the considered baselines, particularly under sustained overload conditions. A multi-seed evaluation with 95% confidence intervals further confirms that the reported gains are robust to stochastic traffic realizations and learning randomness rather than dependent on a single favorable run.