Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)B
Posts
0
Comments
260
Joined
3 yr. ago

  • That seems kind of like pointing to reverse engineering communities and saying that binaries are the preferred format because of how much they can do. Sure you can modify finished models a lot, but what you can do with just pre trained weights vs being able to replicate the final training or changing training parameters is just an entirely different beast.

    There's a reason why the OSI stipulates that code and parameters used to train is considered part of the "source" that should be released in order to count as an open source model.

    You're free to disagree with me and the OSI though, it's not like there's 1 true authority on what open source means. If a game that is highly modifiable and moddable despite the source code not being available counts as open source to you because there are entire communities successfully modding it, then all the more power to you.

  • It's worth noting that OpenR1 have themselves said that DeepSeek didn't release any code for training the models, nor any of the crucial hyperparameters used. So even if you did have suitable training data, you wouldn't be able to replicate it without re-discovering what they did.

    OSI specifically makes a carve-out that allows models to be considered "open source" under their open source AI definition without providing the training data, so when it comes to AI, open source is really about providing the code that kicks off training, checkpoints if used, and details about training data curation so that a comparable dataset can be compiled for replicating the results.

  • It really comes down to this part of the "Open Source" definition:

    The source code [released] must be the preferred form in which a programmer would modify the program

    A compiled binary is not the format in which a programmer would prefer to modify the program - it's much preferred to have the text file which you can edit in a text editor. Just because it's possible to reverse engineer the binary and make changes by patching bytes doesn't make it count. Any programmer would much rather have the source file instead.

    Similarly, the released weights of an AI model are not easy to modify, and are not the "preferred format" that the internal programmers use to make changes to the AI mode. They typically are making changes to the code that does the training and making changes to the training dataset. So for the purpose of calling an AI "open source", the training code and data used to produce the weights are considered the "preferred format", and is what needs to be released for it to really be open source. Internal engineers also typically use training checkpoints, so that they can roll back the model and redo some of the later training steps without redoing all training from the beginning - this is also considered part of the preferred format if it's used.

    OpenR1, which is attempting to recreate R1, notes: No training code was released by DeepSeek, so it is unknown which hyperparameters work best and how they differ across different model families and scales.

    I would call "open weights" models actually just "self hostable" models instead of open source.

  • LDAC

    Jump
  • No problem! You can tell I went deep down the rabbithole a while back lol - I had to rip my dad's CD collection and assure him that what came out of the toslink to his DAC was identical coming from a FLAC as would come from a CD player with optical out.

  • LDAC

    Jump
  • I'm pretty sure if you rip CDs directly to FLAC, it's a perfect copy assuming you're using good software. PCM isn't lossy or lossless because it's not a compressed format, it's an uncompressed bitstream. Think of it like the original data. If it was burned to a CD as digital MP3 data and then ripped that to FLAC, then yes you'd be going from lossy compressed to lossless, which would hide the fact that quality was lost when it went to MP3 in the first place.

    Just as an example, you can rip a CD directly to FLAC (you should also find and use the correct sample offset for your CD drive), rip the cue sheet for track alignment, then burn the FLAC back to a new CD using the cuesheet (and the correct write offset configuration), and you'll get a CD with the exact bit for bit pattern of "pits" burned into the data layer.

    You can then rip both CDs to a raw uncompressed wav file (wav is basically just a container for PCM data) and then you'll be able to MD5sum both wav files and see that they are identical.

    This is how I test my FLAC rips to make sure I'm preserving everything. This is also how CD checksum databases (like CDDB) work - people across the globe can rip to wav or flac and because it's the same master of the CD, they'll get identical checksums, and even after converting the PCM/wav into a flac you are still able to checksum and verify it's identical bit for bit.

  • I think they're making it more complicated than it needs to be. On any other social media site, you find people by their username. So just ask for your friends username (username@instance.com) and put it in the search bar and it'll come up. Using the URL can be convenient on desktop because you can just copy and paste it from the address bar when you're looking at someone's profile.

    And if you want to discover new people where you don't already know their username, then I believe that is the same as any other social media as well, you can come across them in the comments of people you follow or go to the discover tab or search hashtags and you'll find new people that you can tap on and follow.

    I feel like this basically covers how you would find people. A lot of people get hung up on how you know what instance other people are on but it doesn't usually matter. Either someone will give you their username which includes @instance.com, or if you don't know the instance you can search for their name and all known accounts with that name will show up.

    For example if I just search my username "BakedCatboy" (not my real username), the search results show both my mastodon and Pixelfed accounts.

  • It was waitlist for a while, not sure if it still is but I got my welcome email like a week later.

  • Partially yes, the tricky thing is that when using network_mode: "service:tailscale" (presumably on the caddy container since that's what needs to receive traffic from the tailscale network), you won't be able to attach the caddy container to any networks since it's using the tailscale network stack. This means that in order for caddy to reach your containers, you will need to add the tailscale container itself to the relevant networks. Any attached containers will be connected as well.

    (Not sure if I misread the first time or if you edited but the way you say it is right, add the tailscale container to the proxy network so that caddy will also be added and can reach the containers)

    Here's the super condensed version of what matters for connecting traefik/caddy to a VPN like wireguard/tailscale.

    • I left out all WG config since presumably you know how to configure tailscale
    • Left out acme / letsencrypt stuff since that would be different on caddy anyway
    • You may need to configure caddy to trust the tailscale tunnel IP of the machine on the other end that will be reverse proxying over the tunnel.
    • Traefik I guess requires you to specify the docker network to use to reach stuff, I just put anything that should be accessible into "ingress" as you can see. I'm not sure if my setup supports using a different proxy network per app but maybe caddy allows that.

    My traefik compose:

     
        
    services:
      wireguard:
        container_name: wireguard
        networks:
          - ingress
    
      traefik:
        network_mode: "service:wireguard"
        depends_on:
          - wireguard
        command:
          - "--entryPoints.web.proxyProtocol.trustedIPs=10.13.13.1" # Trust remote tunnel IP, the WG container is 10.13.13.2
          - "--entrypoints.websecure.address=:443"
          - "--entryPoints.websecure.proxyProtocol.trustedIPs=10.13.13.1"
          - "--entrypoints.web.http.redirections.entrypoint.to=websecure"
          - "--entrypoints.web.http.redirections.entrypoint.scheme=https"
          - "--entrypoints.web.http.redirections.entrypoint.priority=100"
          - "--providers.docker.exposedByDefault=false"
          - "--providers.docker.network=ingress"
    
    networks:
      ingress:
        external: true
    
    
      

    And then in a service's docker-compose:

     
        
    services:
      ui:
        image: myapp
        read_only: true
        restart: always
        labels:
          - "traefik.enable=true"
          - "traefik.http.routers.myapp.rule=Host(`xxxx.xxxx.xxxx`)"
          - "traefik.http.services.myapp.loadbalancer.server.port=80"
          - "traefik.http.routers.myapp.entrypoints=websecure"
          - "traefik.http.routers.myapp.tls.certresolver=mytlschallenge"
        networks:
          - ingress
    
    networks:
      ingress:
        external: true
    
    
    
      

    (edited to fix formatting on mobile)

  • I've done something similar but I'm not sure how helpful my example would be because I use wireguard instead of tailscale and traefik instead of caddy.

    The principle is the same though, iirc I have my traefik container set to network_mode: "service:wireguard" so that the traefik container uses the wireguard container's network stack. That way the traefik container also sees the wireguard interface and can receive traffic going to the wireguard IP. Then at the other end of the wireguard tunnel I can use haproxy to pass traffic to the wireguard IP through the tunnel and it automatically hits traefik.

  • You could do something like that using point-to-point wireless links or just cables slung between buildings to connect boxes running a self-organizing mesh network protocol like yggdrasil. But there are too many challenges for me to go into depth here ranging from getting buy in from enough people who are located in close proximity, managing user expectations of speed, making services available over such an overlay network (or managing and paying for proxies that provide access to the regular Internet), dealing with geography, etc.

    You'd basically be looking at replicating freifunk or nycmesh or doing something along those lines. NYCmesh as I can tell operates more like an ISP so I would expect it to be at least harder than what they do.

    Imo time is better invested in developing and advancing decentralized applications and protocols, such as developing stuff using bittorrent/DHT or I2P which can just take advantage of the existing internet.

  • I believe 4K is already basically there. I have a 50" 4K (2160p) that I sit 9 feet away from and based on the Nvidia PPD calculator, that makes for 168ppd, and according to that page 150ppd is around the upper limit of human vision. Apple's "retina" displays target around 50-60ppd (varies based on assumed viewing distance), which is what most people seem to consider "average eye visual acuity". Imo 4K / 150ppd is more than enough.

  • If the battery inverter in the Anker box doesn't pass through grid power then I think you would use an automatic transfer switch that switches between mains and battery inverter depending on which is powered. I had dreams of offsetting my homelab power with solar + battery + inverter.

  • Sure no biggie, I keep pretty meticulous records so it's easy to check. My old place in the Boston metro was 4br and used 600-1200kwh, peaking in the summer. Natural gas heat and central AC. Now we're in a 2br in a complex and get more free heat from our neighbors and it ranges from 800-1100, with central heat pump heat and AC, but since the heat isn't gas anymore the heat is included in that.

  • I live in a 2 br apartment and use like 3000-4000 gal per month according to our water bill. Be sure to check if there is a multiplier or if it's in "units of 10" or 100. Our bill used to show multiples of 10 gallons but now it shows in units of 100gal.

  • If the machine doesn't boot then you can use this to access the bios and boot a recovery environment of your choice remotely using pxeboot.

  • Immich has a setting that does automatic photo backup over WiFi, I use the android app as a Google photos replacement. You can choose however many folders on your phone as you want (I just do camera roll) and enable only backup over WiFi and it backs up all the photos in original quality. I self-host the server on my Synology with a reverse proxy (can't forward ports at my current place due to cgnat) so I can access it from anywhere.

    I believe the app is cross platform so the iPhone version should be identical to the android one.

  • Woah federation would be huge!

    Someday I would love to be able to share and receive shared photos / albums to and from users on different servers. Especially if it lets me sync the original files so that I can keep a copy in case their server goes down. It would also be neat if you could enable activitypub so that your account could show up as a fediverse user that people can follow for public or approved follower only posts, pixelfed compatibility would be super cool.

  • Removed Deleted Locked

    Permanently Deleted

    Jump
  • Health insurance execs should live in fear

    😃

    of prison, not murder

    😮‍💨

  • Seconding immich - I host it for my family which makes sharing vacation photos easy since they all have accounts on my instance that can be shared to / from.