Science ❯ Computer Science ❯ Machine Learning
GPT-4 Benchmarking Behavioral Risks ChatGPT-5 Performance Model Optimization Security Vulnerabilities Recursive Self-Improvement Natural Language Processing GPT-5 Performance Consistency End-to-End AI Systems Persistent Models Vulnerability Assessment Claude Mythos Preview Empathy in AI Ethics in AI LLMs Security Applications AI Safety Research and Development Behavioral Patterns Compute Capacity AI Containment Data Management Muse Spark Performance Evaluation OpenAI Self-Exfiltration Llama 4 Human-AI Interaction Hardware Development Testing and Evaluation Claude Reasoning Models Frontier Models AGI and Superintelligence Evaluation of AI Outputs GPT-5.4 AI Decision Making Astra Model Model Training Local AI Deployment Autonomous Systems Claude and GPT Model Training Techniques Vulnerability Research Grok 4 Benchmark Testing Parameter Scaling Gemini Superintelligence World Models Wargaming Simulations GPT-5.4 Benchmarks Performance Metrics Agentic AI Capabilities
The company points to unreleased Astra models that it says met internal research benchmarks after a sandbox escape forced it to pause and tighten training runs.