Showing 2 open source projects for "actor"

View related business solutions
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • 1
    All RL Algorithms from Scratch

    All RL Algorithms from Scratch

    Implementation of all RL algorithms in a simpler way

    ...Its goal is to help learners understand how major reinforcement learning algorithms work under the hood instead of hiding the logic behind large frameworks. The project includes notebooks for value-based methods, policy-gradient methods, actor-critic algorithms, model-based learning, multi-agent reinforcement learning, planning, and hierarchical approaches. Implemented topics include Q-learning, SARSA, Expected SARSA, Dyna-Q, REINFORCE, PPO, A2C, A3C, DDPG, SAC, TRPO, DQN, MADDPG, QMIX, HAC, MCTS, and PlaNet. The code prioritizes clarity, experimentation, and mathematical intuition over production speed. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    MADDPG

    MADDPG

    Code for the MADDPG algorithm from a paper

    MADDPG (Multi-Agent Deep Deterministic Policy Gradient) is the official code release from OpenAI’s paper Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. The repository implements a multi-agent reinforcement learning algorithm that extends DDPG to scenarios where multiple agents interact in shared environments. Each agent has its own policy, but training uses centralized critics conditioned on the observations and actions of all agents, enabling learning in cooperative, competitive, and mixed settings. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next