遇见数据集

Path Mixing VLN dataset

收藏
Zenodo2024-02-12 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

The R2R dataset consists of human annotated instructions corresponding to the paths in these graphs. Each path consists of a sequence of viewpoints encountered by the agent during navigation. A derived dataset, Fine-Grained R2R (FGR2R) [ 12] dataset, annotated parts of instructions with corresponding graph edges to obtain a fine-grained dataset. Existing works in VLN have shown that more instruction examples can improve an agent’s performance in<br> previously unseen environments. Hence to augment training data, we mix parts of paths from the FGR2R dataset to obtain additional instruction-trajectory pairs. The paths are mixed from other neighboring paths that are part of<br> same house which sustains both view and instruction consistency.<br> To mix paths we identify all the edges in the graph corresponding<br> to the start of navigation 𝜀𝑠𝑡𝑎𝑟𝑡 and end of navigational episodes<br> 𝜀𝑒𝑛𝑑 . These edges are important for mixing as they correspond<br> to micro-instructions (Walk away from the desk, Turn right etc.)<br> that refer to start and stop positions in the house while other edges<br> correspond to instructions that back reference to previous locations.<br> The remaining transition edges 𝜀𝑡𝑟𝑎𝑛𝑠 are mixed to obtain a path<br> 𝜀𝑠𝑡𝑎𝑟𝑡 → 𝜀𝑡𝑟𝑎𝑛𝑠 → 𝜀𝑒𝑛𝑑 . Not all edges are inter-connectable, as<br> some of the nodes could be spatially close to each other - reducing<br> the visual variety of viewpoints or resulting in the repetition of<br> micro-instructions (short but actionable instructions) in the final<br> instruction. Accordingly, the edges are connected based on the<br> following criteria: (1) the distance between any 2 nodes should be<br> greater than 3m and the angle between edges should not acute<br> to prevent navigating in loops (2) the distance between the start<br> and end nodes should be greater than 3m to ensure that the path<br> ends up in a different room (3) the start and end nodes cannot<br> have a common edge (4) micro-instructions from common edges<br> of different paths are chosen randomly. The final instruction is<br> the sequence of micro-instructions and the path is the sequence of<br> edges (Figure 2). Using this method, we generate 162k instruction-<br> trajectory pairs with path lengths between 5m and 30m. The final dataset<br> has on average 7.27 views per path, a mean of 14.4m trajectory<br> length and an average of 82 words per instruction.

提供机构:
Zenodo
创建时间:
2023-10-09
二维码
社区交流群
二维码
科研交流群
商业服务