Nestor controlled preview
Start at app.nestor.software. Nestor runs your Inference and Training workloads inside your organization and project. This is the real product; access and workload allowance are enabled by your Nestor preview operator. Contact tirth@nestor.software for activation or help.
1. Get started in the Console
- Sign in using the invited email address and the emailed code.
- Create or select your organization and project. Keep their IDs from Settings for the SDK steps below.
- Ask your Nestor operator to activate the workload allowance for that project. New projects start at zero. An active allowance permits admission checks; it does not promise that capacity is immediately available.
- Choose Serve a model on Overview and follow the Inference path below.
An Application is optional. It gives your AI product a name and a home for linked Endpoints and training runs started from it. It does not deploy a training result automatically, and you can create workloads without creating an Application.
2. Inference: serve, invoke, inspect, delete
For the first preview workload, use this qualified configuration:
| Field | Value |
|---|---|
| Model repository | Qwen/Qwen2.5-0.5B-Instruct |
| Qualified immutable revision | 7ae557604adf67be50417f59c2c2f167def9a775 |
| Runtime, accelerator, region | Auto |
| Replicas | 1 |
Enter the repository in Serve a model. The resolved revision is shown on the Endpoint; check it against the revision above before using this qualification as evidence. If the repository's current revision differs, ask your operator to confirm it. Other models and advanced overrides need separate qualification.
Wait for READY, then send a prompt in the real Playground. Preparation can take several minutes. An accepted Operation is not a ready Endpoint. The detail page shows deployment progress, request IDs, Activity, Metrics and connection information. Playground timings measure the browser request, including network and Console proxy time; Activity and Metrics report server-side measurements.
For an application backend, create an invocation key, preferably bound to this Endpoint, under Access. Use the Endpoint's Connect section for its exact OpenAI-compatible URL. Invocation keys and management API keys have different jobs; see section 4.
When finished, choose Delete. Nestor rejects new invocations as deletion starts. Keep the deletion Operation until cleanup completes; a pending deletion still holds its reservation. Contact your operator if it remains pending for more than ten minutes. Do not create replacements repeatedly while cleanup is unresolved.
3. Training: a finite program and its results
Training is founder-assisted during P1. Arrange a slot with your operator and use the qualified H100, 1 GPU selection. When the available GPU is serving an Endpoint, sequence training after that Endpoint's confirmed deletion and operator sanitation. Do not assume a second simultaneous workload is available merely because the project has allowance for both.
Upload a source snapshot, set the command, and submit. Ask your operator for the reviewed LoRA example source archive: an ordinary Python program with its requirements and a small synthetic support dataset. Repository access is not required. Use:
Command: python -u train_lora.py --steps 8
Mode: Managed Python 3.12
Accelerator: H100
GPUs: 1
Time limit: 1800 seconds (30 minutes)
This example trains small LoRA adapter weights on HuggingFaceTB/SmolLM2-135M
while keeping the base model weights fixed. It proves that a real CUDA program
can train, save a checkpoint, reload its adapter and produce downloadable output.
Eight steps on a tiny synthetic dataset are not evidence of support-answer
quality. These adapters do not apply to the separate Qwen demo Endpoint, and
linking both workloads to an Application does not deploy the adapters there.
Watch the logs, wait for SUCCEEDED, then download the adapter, evaluation file and checkpoint from Files. Record their published SHA-256 hashes and compare downloaded bytes when verifying a result. Outputs must be written to the run's configured output/checkpoint paths; arbitrary working files are not durable results. Canceling requests termination and cleanup; deleting a completed run makes its unpinned files unavailable. Keep the successful run if you want to retain and inspect its results.
4. Access, API keys, SDK and CLI
After the first successful Console workload, have an organization administrator
create a Service account in Access, using the project role your work needs
(ai_admin for approved Inference/Training management). Create its management
API key with an explicit expiry and save the one-time value securely. If issuance
says the organization is not configured, contact your Nestor operator: the
WorkOS organization mapping is an operator prerequisite, not a customer action.
| Credential | Use | Cannot do |
|---|---|---|
| Human Console session | Sign in, manage authorized project resources and issue credentials | Replace an unattended backend credential |
| Service Account management API key | SDK/CLI resource management within its project and role | Invoke an Endpoint or administer IAM/keys |
| Invocation key | Call permitted Endpoints | Create, list, change or delete management resources |
Nestor supplies a versioned preview wheel and SHA-256 checksum from a reviewed
release through a controlled download. The wheel version includes the reviewed
commit, for example 0.1.0+g<12-character-sha>. Install the supplied file, not a
similarly named public package. The durable Python import and CLI are nestor.
# Set this to the actual supplied wheel path, and compare its checksum with the
# release manifest before installing. macOS: shasum -a 256; Linux: sha256sum.
NESTOR_WHEEL=/path/to/reviewed-preview.whl
shasum -a 256 "$NESTOR_WHEEL"
python3 -m venv .venv
. .venv/bin/activate
python -m pip install "$NESTOR_WHEEL"
nestor --version
export NESTOR_API_URL=https://api.nestor.software
export NESTOR_ORGANIZATION=org_your_organization_id
export NESTOR_PROJECT=res_your_project_id
# Bash or zsh: paste the management key at the silent prompt, then press Return.
read -r -s NESTOR_API_KEY
export NESTOR_API_KEY
nestor inference endpoints list
nestor quotas list
Set the API origin and scope explicitly. The current SDK's fallback origin is a
local development address. nestor auth login stores an existing bearer; it is
not a browser sign-in flow and is unnecessary for this machine-key setup.
For a Python invocation, keep the management key above, then supply the separate Endpoint invocation key at the silent prompt:
read -r -s NESTOR_INFERENCE_KEY
export NESTOR_INFERENCE_KEY
export NESTOR_ENDPOINT_ID=endp_your_endpoint_id
python - <<'PY'
import os
from nestor import Nestor
client = Nestor()
endpoint = client.inference.endpoints.get(os.environ["NESTOR_ENDPOINT_ID"])
reply = client.inference.endpoints.chat(
endpoint,
"Give one practical tip for storing an API key safely.",
key=os.environ["NESTOR_INFERENCE_KEY"],
)
print(reply.text)
print("Request:", reply.request_id)
PY
Revoke disposable test keys afterward in Access. For a backend, inject credentials from its secret manager. Never commit them or paste them into support messages. To rotate a management key, create a replacement under the same Service Account, update consumers, verify it works, then revoke the old key. Revoking credentials does not stop already admitted workloads; cancel/delete those explicitly.
5. Troubleshooting and preview limits
| Situation | Next step |
|---|---|
| Activation pending or zero allowance | Give your operator the organization/project IDs and intended workload. |
| Limit already in use | Inspect existing work and pending cleanup; wait or contact the operator. |
| Capacity unavailable | Arrange another slot; repeated submissions do not make capacity available. |
| Endpoint preparing | Follow its deployment stages; allow several minutes. |
| Environment build or program failure | Read build/run logs and exit status; fix dependencies or command before rerunning. |
| Expired/revoked credential | Sign in again for the Console, or replace the appropriate API key. |
| Connection lost during create/delete | Inspect resources and Operations before starting another request; an unchanged idempotent retry can recover the accepted result. |
| Deletion remains pending | Keep the Operation ID and contact the operator after ten minutes. |
| Unexpected platform failure | Send the request/Operation ID and approximate UTC time to the operator. |
Do not include tokens, prompts containing private data, or source secrets in a support report. Preview limits and capacity scheduling are deliberate; there is no billing purchase flow or self-service quota increase. Ask your operator for help with an integration beyond this guide.