Pathway Develops Brain-Inspired BDH Architecture on Amazon SageMaker HyperPod
Pathway is developing BDH, a brain-inspired architecture also known as Dragon Hatchling, on Amazon SageMaker HyperPod with PyTorch integration. The model reasons in latent space rather than generating chain-of-thought tokens, adapting internal memory during inference without test-time weight updates or retraining. SageMaker HyperPod lets Pathway’s scientists share compute as they scale training beyond transformer limits on context windows, forgetting, and costly dense computation.
Pathway is developing its brain-inspired BDH architecture, also called Dragon Hatchling, on Amazon SageMaker HyperPod, integrating the work with well-known frameworks such as PyTorch to scale out training. Instead of externalizing reasoning as a chain-of-thought that generates extra tokens sequentially and feeds them back into later steps, BDH performs reasoning in latent space. It learns from examples and refines a solution without generating an intermediate text trace, moving beyond the transformer paradigm. The architecture was originally formulated as a graph of neurons that communicate through sparse, local interactions and maintain state in synapse-like connections. Model states adapt in context without test-time weight updates, and the reasoning horizon is not limited by a fixed-size context window or tied to a flood of inefficient chain-of-thought tokens. Pathway’s BDH-CQ updates the model’s internal memory during inference by performing iterative computation inside a recurrent latent state and decoding only its candidate answers. Models can work through a new problem without generating long, verbalized reasoning traces, and without requiring fine-tuning or retraining. Large language models tend to forget during long interactions, do not keep knowledge between sessions, and need to be retrained to acquire new knowledge. Amazon SageMaker HyperPod helps Pathway’s applied AI scientists share compute resources in a resilient, scalable, and cost-effective way as they scale out training. Transformers struggle with systematic generalization beyond their training data, particularly for long chain-of-thought reasoning tasks, and require massive amounts of data and computational effort. Dense computation patterns and full back-propagation requirements lead to substantial computational costs that scale exponentially with model size, and it is difficult to efficiently update knowledge without full retraining or fine-tuning. Pathway is using HyperPod and PyTorch to develop BDH as an alternative to that transformer stack.