高级检索

基于深度强化学习与动态掩码的航天器遥感任务调度方法

Spacecraft Remote Sensing Task Scheduling Method Based on Deep Reinforcement Learning and Dynamic Masking

  • 摘要: 针对深空探测领域的航天器遥感任务自主调度问题,提出一种基于深度强化学习(DRL)的序列规划方法。为提高物理建模的精确度,基于分桶查找表计算机动时间,还原探测器在姿态机动过程中的真实动力学过程;针对序列决策时引发的状态漂移,提出一种基于历史决策的时间最近邻状态匹配方法,在解码时逆向检索时间轴上最邻近的已执行节点,作为物理基准来重新校准探测器的当前状态。采用了一维卷积(1D-CNN)与指针网络(Pointer Network)构建端到端的策略网络,在解码端嵌入基于张量计算的动态掩码(Masking),满足多频次观测、时间窗口及能量、存储等约束。数值仿真结果表明,该算法输出的调度序列满足各项物理约束,且在计算时效性、调度收益上具有优势,能够为探测器的自主任务规划提供算法支撑。

     

    Abstract: Aiming at the autonomous scheduling problem of spacecraft remote sensing tasks in the field of deep space exploration, a sequence planning method based on deep reinforcement learning(DRL)is proposed. To improve the accuracy of physical modeling, a bucketized lookup table is used to compute maneuvering time, restoring the real dynamic process during attitude maneuvering. To address state drift caused by sequential decision-making, a time-nearest-neighbor state matching method based on historical decisions is proposed, which reversely retrieves the nearest executed node on the timeline as a physical baseline to recalibrate the current state of the spacecraft. A 1D-CNN and Pointer Network are used to build an end-to-end policy network, with tensor-based dynamic masking embedded at the decoder side to satisfy constraints on multi-frequency observation, time windows, energy, and storage. Simulation results demonstrate that the scheduling sequences produced by the algorithm satisfy all physical constraints and show advantages in computational efficiency and scheduling revenue, providing algorithmic support for autonomous mission planning of spacecraft.

     

/

返回文章
返回