Afhalen na 1 uur in een winkel met voorraad
Gratis thuislevering in België vanaf € 30
Ruim aanbod met 7 miljoen producten

Afhalen na 1 uur in een winkel met voorraad
Gratis thuislevering in België vanaf € 30
Ruim aanbod met 7 miljoen producten

Winkels

Verlanglijstje

Winkels

Verlanglijstje

Zoeken

Volg ons op

140 winkels

Wist je dat er in Vlaanderen voor iedereen een Standaard Boekhandel is binnen een straal van 7 km?

Vind hier een winkel

€ 31,45

+ 62 punten

Levertermijn 1 à 4 weken

Eenvoudig bestellen

Veilig betalen

Gratis thuislevering vanaf € 30 (via bpost)

Gratis levering in je Standaard Boekhandel

Omschrijving

Many real-world problems are inherently hierarchically structured. The use of this structure in an agent's policy may well be the key to improved scalability and higher performance on motor skill tasks. However, such hierarchical structures cannot be exploited by current policy search algorithms. We concentrate on a basic, but highly relevant hierarchy - the `mixed option' policy. Here, a gating network first decides which of the options to execute and, subsequently, the option-policy determines the action. Using a hierarchical setup for our learning method allows us to learn not only one solution to a problem but many. We base our algorithm on a recently proposed information theoretic policy search method, which addresses the exploitation-exploration trade-off by limiting the loss of information between policy updates.